- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明
Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
7.7 KiB
7.7 KiB
项目架构说明
1. 定位
en-news 是一个从英文财经新闻源采集数据,经过正文提取、去重、LLM 翻译/事件抽取、向量化、向量检索,最终生成中文日报的私有化工作流。
设计目标:
- 私有化:不依赖海外服务器,抓取与计算均可在国内 Raspberry Pi 上独立运行。
- 增量友好:每个阶段都尽量跳过已处理文件,断点续跑成本低。
- 幂等:重复运行不会产生重复数据或重复日报。
- 可运维:终端输出阶段/模型/耗时,日志统一写入
logs/。
2. 总体架构
┌─────────────────────────────────────────────────────────────┐
│ en-news 平台 │
│ │
│ configs/ · .env · prompts/ · data/ · logs/ │
│ │
│ app/cli.py ── Typer CLI 入口 │
│ ├── crawl M1 抓取 │
│ ├── extract M2 正文提取 │
│ ├── dedup M3 三层去重 │
│ ├── translate M4 翻译 + 事件抽取 │
│ ├── embed M5 向量生成 │
│ ├── index M6 Qdrant 入库 │
│ ├── search M6 语义搜索 │
│ ├── report M9 日报结构化入库 │
│ ├── pipeline M2→M6(+日报)一键管道 │
│ └── mcp-server M8 MCP 服务 │
│ │
│ scheduler/ ── 管道编排 / 日报生成 │
│ mcp_server/ ── MCP 工具层(研究 Agent 接入) │
└─────────────────────────────────────────────────────────────┘
3. 模块职责
| 模块 | 对应阶段 | 核心职责 |
|---|---|---|
crawler/ |
M1 | 从 RSS/Atom/网页抓取新闻,RSS 优先、Stealth/Headful Web 回退,写入 data/raw/ |
extractor/ |
M2 | 使用 trafilatura(含 Markdown 回退)提取正文,输出 data/processed/ |
dedup/ |
M3 | URL、内容、SimHash 三层去重,SQLite 指纹库,跨源来源合并,输出 data/deduped/ |
llm/ |
M4 | 调用 DeepSeek/Qwen 兼容接口,单次调用完成全文英译中+投资事件抽取,输出 data/events/ |
embedding/ |
M5 | 调用 DashScope text-embedding 批量向量化,输出 data/embeddings/ |
vectorstore/ |
M6 | 管理 Qdrant collection,幂等 upsert、语义搜索、过滤查询 |
scheduler/ |
M7/M9 | 串联 M2→M6,生成每日 AI 摘要日报并写入 MySQL |
mcp_server/ |
M8 | 暴露 MCP 工具供 Cherry Studio/Claude Code 等调用 |
report_db/ |
M9 | MySQL news_report / news_event 建表与读写 |
app/cli.py |
入口 | Typer CLI,各模块的同步调用入口 |
scripts/ |
运维 | Bash 全流程、断点续跑、抓取 Profile、日志清理 |
4. 数据流
M1 抓取
data/raw/{source_id}/{YYYYMMDD}/
├── index.jsonl # 抓取索引(每条含 url_hash、标题、URL)
├── {url_hash}.html # 原始 HTML
└── {url_hash}.md # Crawl4AI 可选的 Markdown
M2 正文提取
data/processed/{source_id}/{YYYYMMDD}/{url_hash}.json
# ProcessedArticle:标题、正文、词数、提取器、时间等
M3 三层去重
data/deduped/{YYYYMMDD}/uniques/{url_hash}.json
# 跨源重复时在已保留唯一篇的 source_ids 中追加来源
M4 翻译 + 事件抽取
data/events/{YYYYMMDD}/{url_hash}.json
# EnTranslatedArticle:中英文标题/正文、事件列表、source_ids
M5 向量生成
data/embeddings/{YYYYMMDD}/{url_hash}.json
# 向量文件(1024 维 text-embedding-v3)+ 每日期 index.json
M6 Qdrant 入库
vectorstore client → collection: en_finance_news
# payload 含标题、双语标题、事件、来源、时间、正文预览
M7/M9 日报
scheduler/reporter.py → MySQL:
news_report(主表:日期、type=intl、AI 摘要、stats)
news_event(明细:Top 事件、来源、情绪、重要度、URL)
5. 关键技术点
5.1 抓取策略
- RSS 优先:大部分源先用 RSS/Atom/Google News RSS,减少对浏览器的依赖。
- Web 回退:RSS 失败/为空时,根据
anti_bot_mode使用 stealth 或 headful Playwright。 - Profile 覆盖:
EN_NEWS_PROFILE环境变量可切换8g_headful/2g_headless等抓取配置。
5.2 三层去重
| 层级 | 维度 | 说明 |
|---|---|---|
| L1 | URL | 相同 URL 直接判重 |
| L2 | 内容 | 正文 hash 相同判重 |
| L3 | SimHash | 汉明距离 ≤ 阈值判为模糊重复 |
指纹库为 data/dedup/fingerprints.sqlite3,SimHash 窗口默认 30 天。
5.3 LLM 场景配置
configs/system.yaml 中通过 llm_scenes 为不同任务配置独立模型:
| 场景 | 用途 | 当前默认 |
|---|---|---|
translation |
M4 翻译+事件抽取 | deepseek-v4-flash,temperature=0.1,max_tokens=8192 |
daily_report |
M7 日报 AI 摘要 | deepseek-v4-flash,temperature=0.3,max_tokens=1500 |
未在场景中声明的字段自动回退到 llm 默认段。Embedding 是单一场景,不能按任务拆分(查询/入库向量必须同模型)。
5.4 增量与幂等
| 阶段 | 增量机制 |
|---|---|
| M2 | processed/{source}/{date}/{url_hash}.json 存在则跳过 |
| M3 | 指纹库判重;重复篇不再写入,只合并来源 |
| M4 | events/{date}/{url_hash}.json 存在则跳过 |
| M5 | embeddings/{date}/{url_hash}.json 存在则跳过 |
| M6 | Qdrant point ID = url_hash 的 UUID,重复 upsert 覆盖 |
| M9 | MySQL 唯一键 (report_date, report_type='intl', file_name='') 覆盖 |
Shell 层还提供步骤级 --resume:状态写在 data/run_state/{YYYYMMDD}.state。
5.5 日报生成模式
- 时间窗口:过去 25 小时(跨天自动聚合)。
- 高重要度事件:
importance >= 4,不足 3 条时逐级降阈。 - AI 摘要:事件超过 10 条自动分批生成,再合并;LLM 失败走规则兜底,不阻断日报入库。
- 日报不再生成 HTML/上传,而是结构化写入 MySQL。
6. 部署拓扑(2026-07 起)
海外服务器(已停用)
↓ 历史:抓取后同步回国内
国内服务器 Pi(当前唯一全链路节点)
/home/pi/intlnews
├── Xvfb :99(headful 浏览器虚拟显示)
├── Privoxy → SOCKS5(HTTP 代理访问海外)
├── crontab: 06:00 / 12:00 / 18:00 / 22:00 执行 domestic_full.sh
├── 本地 Qdrant 文件模式: data/qdrant_storage
└── MySQL: 通过内网隧道访问 news 项目共用 myquant 库
7. 关键目录
app/ CLI 入口
crawler/ M1 抓取
extractor/ M2 正文提取
dedup/ M3 去重
llm/ M4 翻译/事件抽取
embedding/ M5 向量化
vectorstore/ M6 Qdrant
scheduler/ M7/M9 编排与日报
report_db/ M9 MySQL
mcp_server/ M8 MCP
prompts/ LLM Prompt 模板
configs/ 系统/源/Profile 配置
scripts/ Bash 自动化脚本
tests/ 单元测试
docs/ 本文档目录