- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明
Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
3.5 KiB
3.5 KiB
开发指南
1. 代码结构
app/cli.py Typer CLI 入口
crawler/ M1 抓取
config.py 系统配置 + Profile 深度合并
orchestrator.py 串行抓取编排、RSS 优先 Web 回退
rss_crawler.py RSS/Atom/Google News RSS 抓取
crawler.py Crawl4AI Web 抓取
loader.py 读取 sources.yaml
storage.py index.jsonl 存储
extractor/ M2 正文提取
dedup/ M3 三层去重
llm/ M4 LLM 翻译/事件抽取
embedding/ M5 向量化
vectorstore/ M6 Qdrant
scheduler/ M7/M9 管道编排与日报
report_db/ M9 MySQL
mcp_server/ M8 MCP
prompts/ LLM Prompt 模板
scripts/ Bash 自动化与运维
tests/ 单元测试
docs/ 项目文档
2. 常用开发命令
# 安装开发依赖
uv sync --extra dev
# 运行全部测试
uv run pytest
# 运行单个测试文件
uv run pytest tests/test_llm.py
# 代码检查
uv run ruff check .
# 自动修复
uv run ruff check . --fix
# 运行 CLI
uv run en-news --help
3. 测试现状
- 当前测试覆盖:crawler、extractor、dedup、llm、embedding、vectorstore、scheduler、report_db、mcp_server。
- 已知问题:
tests/test_crawler.py::test_write_and_load_index_jsonltests/test_crawler.py::test_write_index_jsonl_dedup- 原因是
crawler/storage.py:load_index()硬编码data/raw/...,未使用测试临时目录。
scheduler/reporter.py与report_db/schema.py因内含长模板/DDL,已在pyproject.toml中忽略 E501。
4. 新增新闻源
- 在
configs/sources.yaml的sources:列表新增条目:- id: "example" name: "Example News" enabled: true homepage: "https://www.example.com/" article_url_pattern: "/news/" js_render: false max_articles_per_run: 20 rss_url: "https://www.example.com/rss" anti_bot_mode: "stealth" - 本地测试:
uv run en-news crawl --source example uv run en-news extract --source example - 确认
data/raw/example/...和data/processed/example/...正常。 - 如果使用 Google News RSS,需确认
rss_crawler.py能正确清理标题和识别跳转链接。
5. LLM Prompt 开发
M4 Prompt 位于 prompts/translation_and_extraction.md。
- 不要随意改变 JSON 输出结构,否则
llm/models.py::LLMTranslationOutput可能校验失败。 - 事件类型定义在
llm/models.py::INTERNATIONAL_EVENT_TYPES,Prompt 中应保持一致。 - 修改后建议跑
uv run pytest tests/test_llm.py。
6. 数据模型核心字段
ProcessedArticle(extractor/models.py):包含source_id、source_ids、标题、URL、正文、词数等。EnTranslatedArticle(llm/models.py):包含中英文标题/正文、events、source_ids。EventExtraction(llm/models.py):事件类型、代码、情绪、重要度、摘要。ReportData/EventRow(report_db/models.py):日报入库结构。
7. 文档维护规范
- 所有面向使用/部署/开发的 Markdown 文档统一放在
docs/。 - 根目录
README.md仅保留仓库入口和文档链接。 - 新增功能或修改流程时,同步更新对应
docs/*.md。 - 不再在仓库根目录维护
CLAUDE.md/continuation.md/english-news-plan.md这类 AI Agent 会话残留。