- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明
Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
105 lines
3.1 KiB
Markdown
105 lines
3.1 KiB
Markdown
# FAQ 常见问题
|
||
|
||
## 1. 为什么有些源只拿到摘要?
|
||
|
||
部分网站存在 Cloudflare / DataDome / 付费墙等反爬限制:
|
||
|
||
- RSS 源只返回标题+摘要(如 Reuters、Investing.com、FT)
|
||
- Web 全文抓取可能被拦截
|
||
- 当前策略是 RSS 优先 + stealth/headful Web 回退,但无法保证 100% 全文
|
||
|
||
解决方法:
|
||
- 使用质量更好的代理/住宅 IP
|
||
- 订阅相应媒体付费内容
|
||
- 对 Pipeline 而言,摘要也能进入翻译/抽取,只是信息量较低
|
||
|
||
## 2. 日报存在哪里?
|
||
|
||
M9 以后日报不再生成 HTML,而是结构化写入 MySQL:
|
||
|
||
- `news_report`:日报主表(`report_type='intl'`)
|
||
- `news_event`:事件明细(`section='intl'`)
|
||
|
||
历史 HTML 日报(2026-06-16 ~ 2026-08-03)仍保留在 `data/reports/` 或已由 news 项目导入。
|
||
|
||
## 3. 如何断点续跑?
|
||
|
||
```bash
|
||
bash scripts/domestic_full.sh --resume
|
||
bash scripts/pipeline.sh --resume
|
||
```
|
||
|
||
- 步骤状态:`data/run_state/{YYYYMMDD}.state`
|
||
- 失败步骤不会标记,`--resume` 会重跑失败步骤
|
||
- 各步骤还有文件级增量,重复运行不会重复处理
|
||
|
||
## 4. 如何添加新源?
|
||
|
||
编辑 `configs/sources.yaml`,参考现有源增加一条配置,然后:
|
||
|
||
```bash
|
||
uv run en-news crawl --source <new_id>
|
||
uv run en-news extract --source <new_id>
|
||
```
|
||
|
||
详细字段见 [配置说明](configuration.md)。
|
||
|
||
## 5. 为什么 translate 或 report 报 API Key/MySQL 错误?
|
||
|
||
检查:
|
||
|
||
```bash
|
||
# 是否已加载 .env
|
||
grep -E 'DEEPSEEK|DASHSCOPE|NEWS_DB' .env | sed 's/=.*/=***/'
|
||
|
||
# DeepSeek
|
||
uv run en-news translate
|
||
|
||
# MySQL
|
||
uv run en-news report
|
||
```
|
||
|
||
CLI 入口会 `load_dotenv()`,Shell 脚本也会加载 `.env`。
|
||
|
||
## 6. Qdrant 本地文件模式和远程模式如何选择?
|
||
|
||
- 默认本地文件模式:`data/qdrant_storage`,零运维,适合 Pi 单机。
|
||
- 远程模式:`.env` 配置非 localhost 的 `QDRANT_URL` 和 `QDRANT_API_KEY`,适合多端共享。
|
||
|
||
切换后需要确保 collection 中向量由同一 embedding 模型生成。
|
||
|
||
## 7. 是否可以只跑管道不生成日报?
|
||
|
||
可以:
|
||
|
||
```bash
|
||
uv run en-news pipeline --skip-report
|
||
```
|
||
|
||
注意 `scripts/pipeline.sh` 目前固定包含日报;如果 MySQL 未配置,可用 CLI `--skip-report` 或改用 `uv run en-news translate`、`embed`、`index` 手动执行。
|
||
|
||
## 8. 为什么日志里出现“AI 大模型(场景 ...)”?
|
||
|
||
这是项目刻意保留的可观测性输出:
|
||
|
||
- M4 会打印 `AI 大模型(场景 translation)`
|
||
- 日报会打印 `AI 大模型(场景 daily_report)`
|
||
- M5 会打印 Embedding 初始化信息
|
||
|
||
方便在日志中确认实际调用的供应商/模型。
|
||
|
||
## 9. 历史 AI Agent 文档去哪了?
|
||
|
||
已清理:
|
||
|
||
- 根目录 `CLAUDE.md`
|
||
- 根目录 `continuation.md`
|
||
- 根目录 `english-news-plan.md`
|
||
- 旧版 `docs/intlnews_usage.html` / `docs/intlnews_usage.md`
|
||
|
||
所有有效内容已整理进当前 `docs/` 文档树。
|
||
|
||
## 10. 测试有失败是否影响生产?
|
||
|
||
当前已知 2 个 `test_crawler.py` 失败,属于测试路径问题,不影响生产管道。建议后续修复 `crawler/storage.py:load_index()` 对测试临时目录的适配。
|