Files
simon 3b44f64f66 refactor: 清理历史 AI Agent 文档残留 + 重构 docs/ + Pipeline 健壮性修复
- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
        统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明

Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
2026-08-22 20:47:53 +08:00

3.1 KiB
Raw Permalink Blame History

FAQ 常见问题

1. 为什么有些源只拿到摘要?

部分网站存在 Cloudflare / DataDome / 付费墙等反爬限制:

  • RSS 源只返回标题+摘要(如 Reuters、Investing.com、FT)
  • Web 全文抓取可能被拦截
  • 当前策略是 RSS 优先 + stealth/headful Web 回退,但无法保证 100% 全文

解决方法:

  • 使用质量更好的代理/住宅 IP
  • 订阅相应媒体付费内容
  • 对 Pipeline 而言,摘要也能进入翻译/抽取,只是信息量较低

2. 日报存在哪里?

M9 以后日报不再生成 HTML,而是结构化写入 MySQL:

  • news_report:日报主表(report_type='intl')
  • news_event:事件明细(section='intl')

历史 HTML 日报(2026-06-16 ~ 2026-08-03)仍保留在 data/reports/ 或已由 news 项目导入。

3. 如何断点续跑?

bash scripts/domestic_full.sh --resume
bash scripts/pipeline.sh --resume
  • 步骤状态:data/run_state/{YYYYMMDD}.state
  • 失败步骤不会标记,--resume 会重跑失败步骤
  • 各步骤还有文件级增量,重复运行不会重复处理

4. 如何添加新源?

编辑 configs/sources.yaml,参考现有源增加一条配置,然后:

uv run en-news crawl --source <new_id>
uv run en-news extract --source <new_id>

详细字段见 配置说明。

5. 为什么 translate 或 report 报 API Key/MySQL 错误?

检查:

# 是否已加载 .env
grep -E 'DEEPSEEK|DASHSCOPE|NEWS_DB' .env | sed 's/=.*/=***/'

# DeepSeek
uv run en-news translate

# MySQL
uv run en-news report

CLI 入口会 load_dotenv(),Shell 脚本也会加载 .env。

6. Qdrant 本地文件模式和远程模式如何选择?

  • 默认本地文件模式:data/qdrant_storage,零运维,适合 Pi 单机。
  • 远程模式:.env 配置非 localhost 的 QDRANT_URL 和 QDRANT_API_KEY,适合多端共享。

切换后需要确保 collection 中向量由同一 embedding 模型生成。

7. 是否可以只跑管道不生成日报?

可以:

uv run en-news pipeline --skip-report

注意 scripts/pipeline.sh 目前固定包含日报;如果 MySQL 未配置,可用 CLI --skip-report 或改用 uv run en-news translate、embed、index 手动执行。

8. 为什么日志里出现“AI 大模型(场景 ...)”?

这是项目刻意保留的可观测性输出:

  • M4 会打印 AI 大模型(场景 translation)
  • 日报会打印 AI 大模型(场景 daily_report)
  • M5 会打印 Embedding 初始化信息

方便在日志中确认实际调用的供应商/模型。

9. 历史 AI Agent 文档去哪了?

已清理:

  • 根目录 CLAUDE.md
  • 根目录 continuation.md
  • 根目录 english-news-plan.md
  • 旧版 docs/intlnews_usage.html / docs/intlnews_usage.md

所有有效内容已整理进当前 docs/ 文档树。

10. 测试有失败是否影响生产?

当前已知 2 个 test_crawler.py 失败,属于测试路径问题,不影响生产管道。建议后续修复 crawler/storage.py:load_index() 对测试临时目录的适配。