refactor: 清理历史 AI Agent 文档残留 + 重构 docs/ + Pipeline 健壮性修复

- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
        统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明

Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
This commit is contained in:
2026-08-22 20:47:53 +08:00
parent 951276a313
commit 3b44f64f66
30 changed files with 1496 additions and 2804 deletions
+104
View File
@@ -0,0 +1,104 @@
# FAQ 常见问题
## 1. 为什么有些源只拿到摘要?
部分网站存在 Cloudflare / DataDome / 付费墙等反爬限制:
- RSS 源只返回标题+摘要(如 Reuters、Investing.com、FT)
- Web 全文抓取可能被拦截
- 当前策略是 RSS 优先 + stealth/headful Web 回退,但无法保证 100% 全文
解决方法:
- 使用质量更好的代理/住宅 IP
- 订阅相应媒体付费内容
- 对 Pipeline 而言,摘要也能进入翻译/抽取,只是信息量较低
## 2. 日报存在哪里?
M9 以后日报不再生成 HTML,而是结构化写入 MySQL:
- `news_report`:日报主表(`report_type='intl'`)
- `news_event`:事件明细(`section='intl'`)
历史 HTML 日报(2026-06-16 ~ 2026-08-03)仍保留在 `data/reports/` 或已由 news 项目导入。
## 3. 如何断点续跑?
```bash
bash scripts/domestic_full.sh --resume
bash scripts/pipeline.sh --resume
```
- 步骤状态:`data/run_state/{YYYYMMDD}.state`
- 失败步骤不会标记,`--resume` 会重跑失败步骤
- 各步骤还有文件级增量,重复运行不会重复处理
## 4. 如何添加新源?
编辑 `configs/sources.yaml`,参考现有源增加一条配置,然后:
```bash
uv run en-news crawl --source <new_id>
uv run en-news extract --source <new_id>
```
详细字段见 [配置说明](configuration.md)。
## 5. 为什么 translate 或 report 报 API Key/MySQL 错误?
检查:
```bash
# 是否已加载 .env
grep -E 'DEEPSEEK|DASHSCOPE|NEWS_DB' .env | sed 's/=.*/=***/'
# DeepSeek
uv run en-news translate
# MySQL
uv run en-news report
```
CLI 入口会 `load_dotenv()`,Shell 脚本也会加载 `.env`。
## 6. Qdrant 本地文件模式和远程模式如何选择?
- 默认本地文件模式:`data/qdrant_storage`,零运维,适合 Pi 单机。
- 远程模式:`.env` 配置非 localhost 的 `QDRANT_URL` 和 `QDRANT_API_KEY`,适合多端共享。
切换后需要确保 collection 中向量由同一 embedding 模型生成。
## 7. 是否可以只跑管道不生成日报?
可以:
```bash
uv run en-news pipeline --skip-report
```
注意 `scripts/pipeline.sh` 目前固定包含日报;如果 MySQL 未配置,可用 CLI `--skip-report` 或改用 `uv run en-news translate`、`embed`、`index` 手动执行。
## 8. 为什么日志里出现“AI 大模型(场景 ...)”?
这是项目刻意保留的可观测性输出:
- M4 会打印 `AI 大模型(场景 translation)`
- 日报会打印 `AI 大模型(场景 daily_report)`
- M5 会打印 Embedding 初始化信息
方便在日志中确认实际调用的供应商/模型。
## 9. 历史 AI Agent 文档去哪了?
已清理:
- 根目录 `CLAUDE.md`
- 根目录 `continuation.md`
- 根目录 `english-news-plan.md`
- 旧版 `docs/intlnews_usage.html` / `docs/intlnews_usage.md`
所有有效内容已整理进当前 `docs/` 文档树。
## 10. 测试有失败是否影响生产?
当前已知 2 个 `test_crawler.py` 失败,属于测试路径问题,不影响生产管道。建议后续修复 `crawler/storage.py:load_index()` 对测试临时目录的适配。