Files
intl_news/docs/architecture.md
T
simon 3b44f64f66 refactor: 清理历史 AI Agent 文档残留 + 重构 docs/ + Pipeline 健壮性修复
- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
        统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明

Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
2026-08-22 20:47:53 +08:00

7.7 KiB
Raw Blame History

项目架构说明

1. 定位

en-news 是一个从英文财经新闻源采集数据,经过正文提取、去重、LLM 翻译/事件抽取、向量化、向量检索,最终生成中文日报的私有化工作流。

设计目标:

  • 私有化:不依赖海外服务器,抓取与计算均可在国内 Raspberry Pi 上独立运行。
  • 增量友好:每个阶段都尽量跳过已处理文件,断点续跑成本低。
  • 幂等:重复运行不会产生重复数据或重复日报。
  • 可运维:终端输出阶段/模型/耗时,日志统一写入 logs/。

2. 总体架构

┌─────────────────────────────────────────────────────────────┐
│                         en-news 平台                          │
│                                                             │
│  configs/  ·  .env  ·  prompts/  ·  data/  ·  logs/         │
│                                                             │
│  app/cli.py ── Typer CLI 入口                                 │
│    ├── crawl      M1 抓取                                     │
│    ├── extract    M2 正文提取                                 │
│    ├── dedup      M3 三层去重                                 │
│    ├── translate  M4 翻译 + 事件抽取                          │
│    ├── embed      M5 向量生成                                 │
│    ├── index      M6 Qdrant 入库                               │
│    ├── search     M6 语义搜索                                 │
│    ├── report     M9 日报结构化入库                            │
│    ├── pipeline   M2→M6(+日报)一键管道                       │
│    └── mcp-server M8 MCP 服务                                 │
│                                                             │
│  scheduler/ ── 管道编排 / 日报生成                            │
│  mcp_server/ ── MCP 工具层(研究 Agent 接入)                 │
└─────────────────────────────────────────────────────────────┘

3. 模块职责

模块 对应阶段 核心职责
crawler/ M1 从 RSS/Atom/网页抓取新闻,RSS 优先、Stealth/Headful Web 回退,写入 data/raw/
extractor/ M2 使用 trafilatura(含 Markdown 回退)提取正文,输出 data/processed/
dedup/ M3 URL、内容、SimHash 三层去重,SQLite 指纹库,跨源来源合并,输出 data/deduped/
llm/ M4 调用 DeepSeek/Qwen 兼容接口,单次调用完成全文英译中+投资事件抽取,输出 data/events/
embedding/ M5 调用 DashScope text-embedding 批量向量化,输出 data/embeddings/
vectorstore/ M6 管理 Qdrant collection,幂等 upsert、语义搜索、过滤查询
scheduler/ M7/M9 串联 M2→M6,生成每日 AI 摘要日报并写入 MySQL
mcp_server/ M8 暴露 MCP 工具供 Cherry Studio/Claude Code 等调用
report_db/ M9 MySQL news_report / news_event 建表与读写
app/cli.py 入口 Typer CLI,各模块的同步调用入口
scripts/ 运维 Bash 全流程、断点续跑、抓取 Profile、日志清理

4. 数据流

M1 抓取
data/raw/{source_id}/{YYYYMMDD}/
  ├── index.jsonl          # 抓取索引(每条含 url_hash、标题、URL)
  ├── {url_hash}.html      # 原始 HTML
  └── {url_hash}.md        # Crawl4AI 可选的 Markdown

M2 正文提取
data/processed/{source_id}/{YYYYMMDD}/{url_hash}.json
  # ProcessedArticle:标题、正文、词数、提取器、时间等

M3 三层去重
data/deduped/{YYYYMMDD}/uniques/{url_hash}.json
  # 跨源重复时在已保留唯一篇的 source_ids 中追加来源

M4 翻译 + 事件抽取
data/events/{YYYYMMDD}/{url_hash}.json
  # EnTranslatedArticle:中英文标题/正文、事件列表、source_ids

M5 向量生成
data/embeddings/{YYYYMMDD}/{url_hash}.json
  # 向量文件(1024 维 text-embedding-v3)+ 每日期 index.json

M6 Qdrant 入库
vectorstore client → collection: en_finance_news
  # payload 含标题、双语标题、事件、来源、时间、正文预览

M7/M9 日报
scheduler/reporter.py → MySQL:
  news_report(主表:日期、type=intl、AI 摘要、stats)
  news_event(明细:Top 事件、来源、情绪、重要度、URL)

5. 关键技术点

5.1 抓取策略

  • RSS 优先:大部分源先用 RSS/Atom/Google News RSS,减少对浏览器的依赖。
  • Web 回退:RSS 失败/为空时,根据 anti_bot_mode 使用 stealth 或 headful Playwright。
  • Profile 覆盖:EN_NEWS_PROFILE 环境变量可切换 8g_headful / 2g_headless 等抓取配置。

5.2 三层去重

层级 维度 说明
L1 URL 相同 URL 直接判重
L2 内容 正文 hash 相同判重
L3 SimHash 汉明距离 ≤ 阈值判为模糊重复

指纹库为 data/dedup/fingerprints.sqlite3,SimHash 窗口默认 30 天。

5.3 LLM 场景配置

configs/system.yaml 中通过 llm_scenes 为不同任务配置独立模型:

场景 用途 当前默认
translation M4 翻译+事件抽取 deepseek-v4-flash,temperature=0.1,max_tokens=8192
daily_report M7 日报 AI 摘要 deepseek-v4-flash,temperature=0.3,max_tokens=1500

未在场景中声明的字段自动回退到 llm 默认段。Embedding 是单一场景,不能按任务拆分(查询/入库向量必须同模型)。

5.4 增量与幂等

阶段 增量机制
M2 processed/{source}/{date}/{url_hash}.json 存在则跳过
M3 指纹库判重;重复篇不再写入,只合并来源
M4 events/{date}/{url_hash}.json 存在则跳过
M5 embeddings/{date}/{url_hash}.json 存在则跳过
M6 Qdrant point ID = url_hash 的 UUID,重复 upsert 覆盖
M9 MySQL 唯一键 (report_date, report_type='intl', file_name='') 覆盖

Shell 层还提供步骤级 --resume:状态写在 data/run_state/{YYYYMMDD}.state。

5.5 日报生成模式

  • 时间窗口:过去 25 小时(跨天自动聚合)。
  • 高重要度事件:importance >= 4,不足 3 条时逐级降阈。
  • AI 摘要:事件超过 10 条自动分批生成,再合并;LLM 失败走规则兜底,不阻断日报入库。
  • 日报不再生成 HTML/上传,而是结构化写入 MySQL。

6. 部署拓扑(2026-07 起)

海外服务器(已停用)
    ↓ 历史:抓取后同步回国内
国内服务器 Pi(当前唯一全链路节点)
  /home/pi/intlnews
    ├── Xvfb :99(headful 浏览器虚拟显示)
    ├── Privoxy → SOCKS5(HTTP 代理访问海外)
    ├── crontab: 06:00 / 12:00 / 18:00 / 22:00 执行 domestic_full.sh
    ├── 本地 Qdrant 文件模式: data/qdrant_storage
    └── MySQL: 通过内网隧道访问 news 项目共用 myquant 库

7. 关键目录

app/           CLI 入口
crawler/       M1 抓取
extractor/     M2 正文提取
dedup/         M3 去重
llm/           M4 翻译/事件抽取
embedding/     M5 向量化
vectorstore/   M6 Qdrant
scheduler/     M7/M9 编排与日报
report_db/     M9 MySQL
mcp_server/    M8 MCP
prompts/       LLM Prompt 模板
configs/       系统/源/Profile 配置
scripts/       Bash 自动化脚本
tests/         单元测试
docs/          本文档目录