Files
intl_news/docs/faq.md
T
simon 3b44f64f66 refactor: 清理历史 AI Agent 文档残留 + 重构 docs/ + Pipeline 健壮性修复
- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
        统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明

Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
2026-08-22 20:47:53 +08:00

105 lines
3.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# FAQ 常见问题
## 1. 为什么有些源只拿到摘要?
部分网站存在 Cloudflare / DataDome / 付费墙等反爬限制:
- RSS 源只返回标题+摘要(如 Reuters、Investing.com、FT)
- Web 全文抓取可能被拦截
- 当前策略是 RSS 优先 + stealth/headful Web 回退,但无法保证 100% 全文
解决方法:
- 使用质量更好的代理/住宅 IP
- 订阅相应媒体付费内容
- 对 Pipeline 而言,摘要也能进入翻译/抽取,只是信息量较低
## 2. 日报存在哪里?
M9 以后日报不再生成 HTML,而是结构化写入 MySQL:
- `news_report`:日报主表(`report_type='intl'`)
- `news_event`:事件明细(`section='intl'`)
历史 HTML 日报(2026-06-16 ~ 2026-08-03)仍保留在 `data/reports/` 或已由 news 项目导入。
## 3. 如何断点续跑?
```bash
bash scripts/domestic_full.sh --resume
bash scripts/pipeline.sh --resume
```
- 步骤状态:`data/run_state/{YYYYMMDD}.state`
- 失败步骤不会标记,`--resume` 会重跑失败步骤
- 各步骤还有文件级增量,重复运行不会重复处理
## 4. 如何添加新源?
编辑 `configs/sources.yaml`,参考现有源增加一条配置,然后:
```bash
uv run en-news crawl --source <new_id>
uv run en-news extract --source <new_id>
```
详细字段见 [配置说明](configuration.md)。
## 5. 为什么 translate 或 report 报 API Key/MySQL 错误?
检查:
```bash
# 是否已加载 .env
grep -E 'DEEPSEEK|DASHSCOPE|NEWS_DB' .env | sed 's/=.*/=***/'
# DeepSeek
uv run en-news translate
# MySQL
uv run en-news report
```
CLI 入口会 `load_dotenv()`,Shell 脚本也会加载 `.env`。
## 6. Qdrant 本地文件模式和远程模式如何选择?
- 默认本地文件模式:`data/qdrant_storage`,零运维,适合 Pi 单机。
- 远程模式:`.env` 配置非 localhost 的 `QDRANT_URL` 和 `QDRANT_API_KEY`,适合多端共享。
切换后需要确保 collection 中向量由同一 embedding 模型生成。
## 7. 是否可以只跑管道不生成日报?
可以:
```bash
uv run en-news pipeline --skip-report
```
注意 `scripts/pipeline.sh` 目前固定包含日报;如果 MySQL 未配置,可用 CLI `--skip-report` 或改用 `uv run en-news translate`、`embed`、`index` 手动执行。
## 8. 为什么日志里出现“AI 大模型(场景 ...)”?
这是项目刻意保留的可观测性输出:
- M4 会打印 `AI 大模型(场景 translation)`
- 日报会打印 `AI 大模型(场景 daily_report)`
- M5 会打印 Embedding 初始化信息
方便在日志中确认实际调用的供应商/模型。
## 9. 历史 AI Agent 文档去哪了?
已清理:
- 根目录 `CLAUDE.md`
- 根目录 `continuation.md`
- 根目录 `english-news-plan.md`
- 旧版 `docs/intlnews_usage.html` / `docs/intlnews_usage.md`
所有有效内容已整理进当前 `docs/` 文档树。
## 10. 测试有失败是否影响生产?
当前已知 2 个 `test_crawler.py` 失败,属于测试路径问题,不影响生产管道。建议后续修复 `crawler/storage.py:load_index()` 对测试临时目录的适配。