- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明
Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
3.1 KiB
3.1 KiB
FAQ 常见问题
1. 为什么有些源只拿到摘要?
部分网站存在 Cloudflare / DataDome / 付费墙等反爬限制:
- RSS 源只返回标题+摘要(如 Reuters、Investing.com、FT)
- Web 全文抓取可能被拦截
- 当前策略是 RSS 优先 + stealth/headful Web 回退,但无法保证 100% 全文
解决方法:
- 使用质量更好的代理/住宅 IP
- 订阅相应媒体付费内容
- 对 Pipeline 而言,摘要也能进入翻译/抽取,只是信息量较低
2. 日报存在哪里?
M9 以后日报不再生成 HTML,而是结构化写入 MySQL:
news_report:日报主表(report_type='intl')news_event:事件明细(section='intl')
历史 HTML 日报(2026-06-16 ~ 2026-08-03)仍保留在 data/reports/ 或已由 news 项目导入。
3. 如何断点续跑?
bash scripts/domestic_full.sh --resume
bash scripts/pipeline.sh --resume
- 步骤状态:
data/run_state/{YYYYMMDD}.state - 失败步骤不会标记,
--resume会重跑失败步骤 - 各步骤还有文件级增量,重复运行不会重复处理
4. 如何添加新源?
编辑 configs/sources.yaml,参考现有源增加一条配置,然后:
uv run en-news crawl --source <new_id>
uv run en-news extract --source <new_id>
详细字段见 配置说明。
5. 为什么 translate 或 report 报 API Key/MySQL 错误?
检查:
# 是否已加载 .env
grep -E 'DEEPSEEK|DASHSCOPE|NEWS_DB' .env | sed 's/=.*/=***/'
# DeepSeek
uv run en-news translate
# MySQL
uv run en-news report
CLI 入口会 load_dotenv(),Shell 脚本也会加载 .env。
6. Qdrant 本地文件模式和远程模式如何选择?
- 默认本地文件模式:
data/qdrant_storage,零运维,适合 Pi 单机。 - 远程模式:
.env配置非 localhost 的QDRANT_URL和QDRANT_API_KEY,适合多端共享。
切换后需要确保 collection 中向量由同一 embedding 模型生成。
7. 是否可以只跑管道不生成日报?
可以:
uv run en-news pipeline --skip-report
注意 scripts/pipeline.sh 目前固定包含日报;如果 MySQL 未配置,可用 CLI --skip-report 或改用 uv run en-news translate、embed、index 手动执行。
8. 为什么日志里出现“AI 大模型(场景 ...)”?
这是项目刻意保留的可观测性输出:
- M4 会打印
AI 大模型(场景 translation) - 日报会打印
AI 大模型(场景 daily_report) - M5 会打印 Embedding 初始化信息
方便在日志中确认实际调用的供应商/模型。
9. 历史 AI Agent 文档去哪了?
已清理:
- 根目录
CLAUDE.md - 根目录
continuation.md - 根目录
english-news-plan.md - 旧版
docs/intlnews_usage.html/docs/intlnews_usage.md
所有有效内容已整理进当前 docs/ 文档树。
10. 测试有失败是否影响生产?
当前已知 2 个 test_crawler.py 失败,属于测试路径问题,不影响生产管道。建议后续修复 crawler/storage.py:load_index() 对测试临时目录的适配。