refactor: 清理历史 AI Agent 文档残留 + 重构 docs/ + Pipeline 健壮性修复

- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
        统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明

Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
This commit is contained in:
2026-08-22 20:47:53 +08:00
parent 951276a313
commit 3b44f64f66
30 changed files with 1496 additions and 2804 deletions
+101
View File
@@ -0,0 +1,101 @@
# 开发指南
## 1. 代码结构
```text
app/cli.py Typer CLI 入口
crawler/ M1 抓取
config.py 系统配置 + Profile 深度合并
orchestrator.py 串行抓取编排、RSS 优先 Web 回退
rss_crawler.py RSS/Atom/Google News RSS 抓取
crawler.py Crawl4AI Web 抓取
loader.py 读取 sources.yaml
storage.py index.jsonl 存储
extractor/ M2 正文提取
dedup/ M3 三层去重
llm/ M4 LLM 翻译/事件抽取
embedding/ M5 向量化
vectorstore/ M6 Qdrant
scheduler/ M7/M9 管道编排与日报
report_db/ M9 MySQL
mcp_server/ M8 MCP
prompts/ LLM Prompt 模板
scripts/ Bash 自动化与运维
tests/ 单元测试
docs/ 项目文档
```
## 2. 常用开发命令
```bash
# 安装开发依赖
uv sync --extra dev
# 运行全部测试
uv run pytest
# 运行单个测试文件
uv run pytest tests/test_llm.py
# 代码检查
uv run ruff check .
# 自动修复
uv run ruff check . --fix
# 运行 CLI
uv run en-news --help
```
## 3. 测试现状
- 当前测试覆盖:crawler、extractor、dedup、llm、embedding、vectorstore、scheduler、report_db、mcp_server。
- 已知问题:
- `tests/test_crawler.py::test_write_and_load_index_jsonl`
- `tests/test_crawler.py::test_write_index_jsonl_dedup`
- 原因是 `crawler/storage.py:load_index()` 硬编码 `data/raw/...`,未使用测试临时目录。
- `scheduler/reporter.py` 与 `report_db/schema.py` 因内含长模板/DDL,已在 `pyproject.toml` 中忽略 E501。
## 4. 新增新闻源
1. 在 `configs/sources.yaml` 的 `sources:` 列表新增条目:
```yaml
- id: "example"
name: "Example News"
enabled: true
homepage: "https://www.example.com/"
article_url_pattern: "/news/"
js_render: false
max_articles_per_run: 20
rss_url: "https://www.example.com/rss"
anti_bot_mode: "stealth"
```
2. 本地测试:
```bash
uv run en-news crawl --source example
uv run en-news extract --source example
```
3. 确认 `data/raw/example/...` 和 `data/processed/example/...` 正常。
4. 如果使用 Google News RSS,需确认 `rss_crawler.py` 能正确清理标题和识别跳转链接。
## 5. LLM Prompt 开发
M4 Prompt 位于 `prompts/translation_and_extraction.md`。
- 不要随意改变 JSON 输出结构,否则 `llm/models.py::LLMTranslationOutput` 可能校验失败。
- 事件类型定义在 `llm/models.py::INTERNATIONAL_EVENT_TYPES`,Prompt 中应保持一致。
- 修改后建议跑 `uv run pytest tests/test_llm.py`。
## 6. 数据模型核心字段
- `ProcessedArticle`(`extractor/models.py`):包含 `source_id`、`source_ids`、标题、URL、正文、词数等。
- `EnTranslatedArticle`(`llm/models.py`):包含中英文标题/正文、`events`、`source_ids`。
- `EventExtraction`(`llm/models.py`):事件类型、代码、情绪、重要度、摘要。
- `ReportData` / `EventRow`(`report_db/models.py`):日报入库结构。
## 7. 文档维护规范
- 所有面向使用/部署/开发的 Markdown 文档统一放在 `docs/`。
- 根目录 `README.md` 仅保留仓库入口和文档链接。
- 新增功能或修改流程时,同步更新对应 `docs/*.md`。
- 不再在仓库根目录维护 `CLAUDE.md` / `continuation.md` / `english-news-plan.md` 这类 AI Agent 会话残留。