Files
simon 3b44f64f66 refactor: 清理历史 AI Agent 文档残留 + 重构 docs/ + Pipeline 健壮性修复
- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
        统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明

Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
2026-08-22 20:47:53 +08:00

3.5 KiB
Raw Permalink Blame History

开发指南

1. 代码结构

app/cli.py              Typer CLI 入口
crawler/                M1 抓取
  config.py             系统配置 + Profile 深度合并
  orchestrator.py       串行抓取编排、RSS 优先 Web 回退
  rss_crawler.py        RSS/Atom/Google News RSS 抓取
  crawler.py            Crawl4AI Web 抓取
  loader.py             读取 sources.yaml
  storage.py            index.jsonl 存储
extractor/              M2 正文提取
dedup/                  M3 三层去重
llm/                    M4 LLM 翻译/事件抽取
embedding/              M5 向量化
vectorstore/            M6 Qdrant
scheduler/              M7/M9 管道编排与日报
report_db/              M9 MySQL
mcp_server/             M8 MCP
prompts/                LLM Prompt 模板
scripts/                Bash 自动化与运维
tests/                  单元测试
docs/                   项目文档

2. 常用开发命令

# 安装开发依赖
uv sync --extra dev

# 运行全部测试
uv run pytest

# 运行单个测试文件
uv run pytest tests/test_llm.py

# 代码检查
uv run ruff check .

# 自动修复
uv run ruff check . --fix

# 运行 CLI
uv run en-news --help

3. 测试现状

  • 当前测试覆盖:crawler、extractor、dedup、llm、embedding、vectorstore、scheduler、report_db、mcp_server。
  • 已知问题:
    • tests/test_crawler.py::test_write_and_load_index_jsonl
    • tests/test_crawler.py::test_write_index_jsonl_dedup
    • 原因是 crawler/storage.py:load_index() 硬编码 data/raw/...,未使用测试临时目录。
  • scheduler/reporter.py 与 report_db/schema.py 因内含长模板/DDL,已在 pyproject.toml 中忽略 E501。

4. 新增新闻源

  1. 在 configs/sources.yaml 的 sources: 列表新增条目:
    - id: "example"
      name: "Example News"
      enabled: true
      homepage: "https://www.example.com/"
      article_url_pattern: "/news/"
      js_render: false
      max_articles_per_run: 20
      rss_url: "https://www.example.com/rss"
      anti_bot_mode: "stealth"
    
  2. 本地测试:
    uv run en-news crawl --source example
    uv run en-news extract --source example
    
  3. 确认 data/raw/example/... 和 data/processed/example/... 正常。
  4. 如果使用 Google News RSS,需确认 rss_crawler.py 能正确清理标题和识别跳转链接。

5. LLM Prompt 开发

M4 Prompt 位于 prompts/translation_and_extraction.md。

  • 不要随意改变 JSON 输出结构,否则 llm/models.py::LLMTranslationOutput 可能校验失败。
  • 事件类型定义在 llm/models.py::INTERNATIONAL_EVENT_TYPES,Prompt 中应保持一致。
  • 修改后建议跑 uv run pytest tests/test_llm.py。

6. 数据模型核心字段

  • ProcessedArticle(extractor/models.py):包含 source_id、source_ids、标题、URL、正文、词数等。
  • EnTranslatedArticle(llm/models.py):包含中英文标题/正文、events、source_ids。
  • EventExtraction(llm/models.py):事件类型、代码、情绪、重要度、摘要。
  • ReportData / EventRow(report_db/models.py):日报入库结构。

7. 文档维护规范

  • 所有面向使用/部署/开发的 Markdown 文档统一放在 docs/。
  • 根目录 README.md 仅保留仓库入口和文档链接。
  • 新增功能或修改流程时,同步更新对应 docs/*.md。
  • 不再在仓库根目录维护 CLAUDE.md / continuation.md / english-news-plan.md 这类 AI Agent 会话残留。