Files
intl_news/docs/development.md
T
simon 3b44f64f66 refactor: 清理历史 AI Agent 文档残留 + 重构 docs/ + Pipeline 健壮性修复
- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
        统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明

Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
2026-08-22 20:47:53 +08:00

102 lines
3.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 开发指南
## 1. 代码结构
```text
app/cli.py Typer CLI 入口
crawler/ M1 抓取
config.py 系统配置 + Profile 深度合并
orchestrator.py 串行抓取编排、RSS 优先 Web 回退
rss_crawler.py RSS/Atom/Google News RSS 抓取
crawler.py Crawl4AI Web 抓取
loader.py 读取 sources.yaml
storage.py index.jsonl 存储
extractor/ M2 正文提取
dedup/ M3 三层去重
llm/ M4 LLM 翻译/事件抽取
embedding/ M5 向量化
vectorstore/ M6 Qdrant
scheduler/ M7/M9 管道编排与日报
report_db/ M9 MySQL
mcp_server/ M8 MCP
prompts/ LLM Prompt 模板
scripts/ Bash 自动化与运维
tests/ 单元测试
docs/ 项目文档
```
## 2. 常用开发命令
```bash
# 安装开发依赖
uv sync --extra dev
# 运行全部测试
uv run pytest
# 运行单个测试文件
uv run pytest tests/test_llm.py
# 代码检查
uv run ruff check .
# 自动修复
uv run ruff check . --fix
# 运行 CLI
uv run en-news --help
```
## 3. 测试现状
- 当前测试覆盖:crawler、extractor、dedup、llm、embedding、vectorstore、scheduler、report_db、mcp_server。
- 已知问题:
- `tests/test_crawler.py::test_write_and_load_index_jsonl`
- `tests/test_crawler.py::test_write_index_jsonl_dedup`
- 原因是 `crawler/storage.py:load_index()` 硬编码 `data/raw/...`,未使用测试临时目录。
- `scheduler/reporter.py` 与 `report_db/schema.py` 因内含长模板/DDL,已在 `pyproject.toml` 中忽略 E501。
## 4. 新增新闻源
1. 在 `configs/sources.yaml` 的 `sources:` 列表新增条目:
```yaml
- id: "example"
name: "Example News"
enabled: true
homepage: "https://www.example.com/"
article_url_pattern: "/news/"
js_render: false
max_articles_per_run: 20
rss_url: "https://www.example.com/rss"
anti_bot_mode: "stealth"
```
2. 本地测试:
```bash
uv run en-news crawl --source example
uv run en-news extract --source example
```
3. 确认 `data/raw/example/...` 和 `data/processed/example/...` 正常。
4. 如果使用 Google News RSS,需确认 `rss_crawler.py` 能正确清理标题和识别跳转链接。
## 5. LLM Prompt 开发
M4 Prompt 位于 `prompts/translation_and_extraction.md`。
- 不要随意改变 JSON 输出结构,否则 `llm/models.py::LLMTranslationOutput` 可能校验失败。
- 事件类型定义在 `llm/models.py::INTERNATIONAL_EVENT_TYPES`,Prompt 中应保持一致。
- 修改后建议跑 `uv run pytest tests/test_llm.py`。
## 6. 数据模型核心字段
- `ProcessedArticle`(`extractor/models.py`):包含 `source_id`、`source_ids`、标题、URL、正文、词数等。
- `EnTranslatedArticle`(`llm/models.py`):包含中英文标题/正文、`events`、`source_ids`。
- `EventExtraction`(`llm/models.py`):事件类型、代码、情绪、重要度、摘要。
- `ReportData` / `EventRow`(`report_db/models.py`):日报入库结构。
## 7. 文档维护规范
- 所有面向使用/部署/开发的 Markdown 文档统一放在 `docs/`。
- 根目录 `README.md` 仅保留仓库入口和文档链接。
- 新增功能或修改流程时,同步更新对应 `docs/*.md`。
- 不再在仓库根目录维护 `CLAUDE.md` / `continuation.md` / `english-news-plan.md` 这类 AI Agent 会话残留。