# 开发指南 ## 1. 代码结构 ```text app/cli.py Typer CLI 入口 crawler/ M1 抓取 config.py 系统配置 + Profile 深度合并 orchestrator.py 串行抓取编排、RSS 优先 Web 回退 rss_crawler.py RSS/Atom/Google News RSS 抓取 crawler.py Crawl4AI Web 抓取 loader.py 读取 sources.yaml storage.py index.jsonl 存储 extractor/ M2 正文提取 dedup/ M3 三层去重 llm/ M4 LLM 翻译/事件抽取 embedding/ M5 向量化 vectorstore/ M6 Qdrant scheduler/ M7/M9 管道编排与日报 report_db/ M9 MySQL mcp_server/ M8 MCP prompts/ LLM Prompt 模板 scripts/ Bash 自动化与运维 tests/ 单元测试 docs/ 项目文档 ``` ## 2. 常用开发命令 ```bash # 安装开发依赖 uv sync --extra dev # 运行全部测试 uv run pytest # 运行单个测试文件 uv run pytest tests/test_llm.py # 代码检查 uv run ruff check . # 自动修复 uv run ruff check . --fix # 运行 CLI uv run en-news --help ``` ## 3. 测试现状 - 当前测试覆盖:crawler、extractor、dedup、llm、embedding、vectorstore、scheduler、report_db、mcp_server。 - 已知问题: - `tests/test_crawler.py::test_write_and_load_index_jsonl` - `tests/test_crawler.py::test_write_index_jsonl_dedup` - 原因是 `crawler/storage.py:load_index()` 硬编码 `data/raw/...`,未使用测试临时目录。 - `scheduler/reporter.py` 与 `report_db/schema.py` 因内含长模板/DDL,已在 `pyproject.toml` 中忽略 E501。 ## 4. 新增新闻源 1. 在 `configs/sources.yaml` 的 `sources:` 列表新增条目: ```yaml - id: "example" name: "Example News" enabled: true homepage: "https://www.example.com/" article_url_pattern: "/news/" js_render: false max_articles_per_run: 20 rss_url: "https://www.example.com/rss" anti_bot_mode: "stealth" ``` 2. 本地测试: ```bash uv run en-news crawl --source example uv run en-news extract --source example ``` 3. 确认 `data/raw/example/...` 和 `data/processed/example/...` 正常。 4. 如果使用 Google News RSS,需确认 `rss_crawler.py` 能正确清理标题和识别跳转链接。 ## 5. LLM Prompt 开发 M4 Prompt 位于 `prompts/translation_and_extraction.md`。 - 不要随意改变 JSON 输出结构,否则 `llm/models.py::LLMTranslationOutput` 可能校验失败。 - 事件类型定义在 `llm/models.py::INTERNATIONAL_EVENT_TYPES`,Prompt 中应保持一致。 - 修改后建议跑 `uv run pytest tests/test_llm.py`。 ## 6. 数据模型核心字段 - `ProcessedArticle`(`extractor/models.py`):包含 `source_id`、`source_ids`、标题、URL、正文、词数等。 - `EnTranslatedArticle`(`llm/models.py`):包含中英文标题/正文、`events`、`source_ids`。 - `EventExtraction`(`llm/models.py`):事件类型、代码、情绪、重要度、摘要。 - `ReportData` / `EventRow`(`report_db/models.py`):日报入库结构。 ## 7. 文档维护规范 - 所有面向使用/部署/开发的 Markdown 文档统一放在 `docs/`。 - 根目录 `README.md` 仅保留仓库入口和文档链接。 - 新增功能或修改流程时,同步更新对应 `docs/*.md`。 - 不再在仓库根目录维护 `CLAUDE.md` / `continuation.md` / `english-news-plan.md` 这类 AI Agent 会话残留。