- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明
Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
102 lines
3.5 KiB
Markdown
102 lines
3.5 KiB
Markdown
# 开发指南
|
||
|
||
## 1. 代码结构
|
||
|
||
```text
|
||
app/cli.py Typer CLI 入口
|
||
crawler/ M1 抓取
|
||
config.py 系统配置 + Profile 深度合并
|
||
orchestrator.py 串行抓取编排、RSS 优先 Web 回退
|
||
rss_crawler.py RSS/Atom/Google News RSS 抓取
|
||
crawler.py Crawl4AI Web 抓取
|
||
loader.py 读取 sources.yaml
|
||
storage.py index.jsonl 存储
|
||
extractor/ M2 正文提取
|
||
dedup/ M3 三层去重
|
||
llm/ M4 LLM 翻译/事件抽取
|
||
embedding/ M5 向量化
|
||
vectorstore/ M6 Qdrant
|
||
scheduler/ M7/M9 管道编排与日报
|
||
report_db/ M9 MySQL
|
||
mcp_server/ M8 MCP
|
||
prompts/ LLM Prompt 模板
|
||
scripts/ Bash 自动化与运维
|
||
tests/ 单元测试
|
||
docs/ 项目文档
|
||
```
|
||
|
||
## 2. 常用开发命令
|
||
|
||
```bash
|
||
# 安装开发依赖
|
||
uv sync --extra dev
|
||
|
||
# 运行全部测试
|
||
uv run pytest
|
||
|
||
# 运行单个测试文件
|
||
uv run pytest tests/test_llm.py
|
||
|
||
# 代码检查
|
||
uv run ruff check .
|
||
|
||
# 自动修复
|
||
uv run ruff check . --fix
|
||
|
||
# 运行 CLI
|
||
uv run en-news --help
|
||
```
|
||
|
||
## 3. 测试现状
|
||
|
||
- 当前测试覆盖:crawler、extractor、dedup、llm、embedding、vectorstore、scheduler、report_db、mcp_server。
|
||
- 已知问题:
|
||
- `tests/test_crawler.py::test_write_and_load_index_jsonl`
|
||
- `tests/test_crawler.py::test_write_index_jsonl_dedup`
|
||
- 原因是 `crawler/storage.py:load_index()` 硬编码 `data/raw/...`,未使用测试临时目录。
|
||
- `scheduler/reporter.py` 与 `report_db/schema.py` 因内含长模板/DDL,已在 `pyproject.toml` 中忽略 E501。
|
||
|
||
## 4. 新增新闻源
|
||
|
||
1. 在 `configs/sources.yaml` 的 `sources:` 列表新增条目:
|
||
```yaml
|
||
- id: "example"
|
||
name: "Example News"
|
||
enabled: true
|
||
homepage: "https://www.example.com/"
|
||
article_url_pattern: "/news/"
|
||
js_render: false
|
||
max_articles_per_run: 20
|
||
rss_url: "https://www.example.com/rss"
|
||
anti_bot_mode: "stealth"
|
||
```
|
||
2. 本地测试:
|
||
```bash
|
||
uv run en-news crawl --source example
|
||
uv run en-news extract --source example
|
||
```
|
||
3. 确认 `data/raw/example/...` 和 `data/processed/example/...` 正常。
|
||
4. 如果使用 Google News RSS,需确认 `rss_crawler.py` 能正确清理标题和识别跳转链接。
|
||
|
||
## 5. LLM Prompt 开发
|
||
|
||
M4 Prompt 位于 `prompts/translation_and_extraction.md`。
|
||
|
||
- 不要随意改变 JSON 输出结构,否则 `llm/models.py::LLMTranslationOutput` 可能校验失败。
|
||
- 事件类型定义在 `llm/models.py::INTERNATIONAL_EVENT_TYPES`,Prompt 中应保持一致。
|
||
- 修改后建议跑 `uv run pytest tests/test_llm.py`。
|
||
|
||
## 6. 数据模型核心字段
|
||
|
||
- `ProcessedArticle`(`extractor/models.py`):包含 `source_id`、`source_ids`、标题、URL、正文、词数等。
|
||
- `EnTranslatedArticle`(`llm/models.py`):包含中英文标题/正文、`events`、`source_ids`。
|
||
- `EventExtraction`(`llm/models.py`):事件类型、代码、情绪、重要度、摘要。
|
||
- `ReportData` / `EventRow`(`report_db/models.py`):日报入库结构。
|
||
|
||
## 7. 文档维护规范
|
||
|
||
- 所有面向使用/部署/开发的 Markdown 文档统一放在 `docs/`。
|
||
- 根目录 `README.md` 仅保留仓库入口和文档链接。
|
||
- 新增功能或修改流程时,同步更新对应 `docs/*.md`。
|
||
- 不再在仓库根目录维护 `CLAUDE.md` / `continuation.md` / `english-news-plan.md` 这类 AI Agent 会话残留。
|