# 项目架构说明 ## 1. 定位 `en-news` 是一个从英文财经新闻源采集数据,经过正文提取、去重、LLM 翻译/事件抽取、向量化、向量检索,最终生成中文日报的私有化工作流。 设计目标: - **私有化**:不依赖海外服务器,抓取与计算均可在国内 Raspberry Pi 上独立运行。 - **增量友好**:每个阶段都尽量跳过已处理文件,断点续跑成本低。 - **幂等**:重复运行不会产生重复数据或重复日报。 - **可运维**:终端输出阶段/模型/耗时,日志统一写入 `logs/`。 ## 2. 总体架构 ```text ┌─────────────────────────────────────────────────────────────┐ │ en-news 平台 │ │ │ │ configs/ · .env · prompts/ · data/ · logs/ │ │ │ │ app/cli.py ── Typer CLI 入口 │ │ ├── crawl M1 抓取 │ │ ├── extract M2 正文提取 │ │ ├── dedup M3 三层去重 │ │ ├── translate M4 翻译 + 事件抽取 │ │ ├── embed M5 向量生成 │ │ ├── index M6 Qdrant 入库 │ │ ├── search M6 语义搜索 │ │ ├── report M9 日报结构化入库 │ │ ├── pipeline M2→M6(+日报)一键管道 │ │ └── mcp-server M8 MCP 服务 │ │ │ │ scheduler/ ── 管道编排 / 日报生成 │ │ mcp_server/ ── MCP 工具层(研究 Agent 接入) │ └─────────────────────────────────────────────────────────────┘ ``` ## 3. 模块职责 | 模块 | 对应阶段 | 核心职责 | |------|----------|----------| | `crawler/` | M1 | 从 RSS/Atom/网页抓取新闻,RSS 优先、Stealth/Headful Web 回退,写入 `data/raw/` | | `extractor/` | M2 | 使用 trafilatura(含 Markdown 回退)提取正文,输出 `data/processed/` | | `dedup/` | M3 | URL、内容、SimHash 三层去重,SQLite 指纹库,跨源来源合并,输出 `data/deduped/` | | `llm/` | M4 | 调用 DeepSeek/Qwen 兼容接口,单次调用完成全文英译中+投资事件抽取,输出 `data/events/` | | `embedding/` | M5 | 调用 DashScope text-embedding 批量向量化,输出 `data/embeddings/` | | `vectorstore/` | M6 | 管理 Qdrant collection,幂等 upsert、语义搜索、过滤查询 | | `scheduler/` | M7/M9 | 串联 M2→M6,生成每日 AI 摘要日报并写入 MySQL | | `mcp_server/` | M8 | 暴露 MCP 工具供 Cherry Studio/Claude Code 等调用 | | `report_db/` | M9 | MySQL `news_report` / `news_event` 建表与读写 | | `app/cli.py` | 入口 | Typer CLI,各模块的同步调用入口 | | `scripts/` | 运维 | Bash 全流程、断点续跑、抓取 Profile、日志清理 | ## 4. 数据流 ```text M1 抓取 data/raw/{source_id}/{YYYYMMDD}/ ├── index.jsonl # 抓取索引(每条含 url_hash、标题、URL) ├── {url_hash}.html # 原始 HTML └── {url_hash}.md # Crawl4AI 可选的 Markdown M2 正文提取 data/processed/{source_id}/{YYYYMMDD}/{url_hash}.json # ProcessedArticle:标题、正文、词数、提取器、时间等 M3 三层去重 data/deduped/{YYYYMMDD}/uniques/{url_hash}.json # 跨源重复时在已保留唯一篇的 source_ids 中追加来源 M4 翻译 + 事件抽取 data/events/{YYYYMMDD}/{url_hash}.json # EnTranslatedArticle:中英文标题/正文、事件列表、source_ids M5 向量生成 data/embeddings/{YYYYMMDD}/{url_hash}.json # 向量文件(1024 维 text-embedding-v3)+ 每日期 index.json M6 Qdrant 入库 vectorstore client → collection: en_finance_news # payload 含标题、双语标题、事件、来源、时间、正文预览 M7/M9 日报 scheduler/reporter.py → MySQL: news_report(主表:日期、type=intl、AI 摘要、stats) news_event(明细:Top 事件、来源、情绪、重要度、URL) ``` ## 5. 关键技术点 ### 5.1 抓取策略 - **RSS 优先**:大部分源先用 RSS/Atom/Google News RSS,减少对浏览器的依赖。 - **Web 回退**:RSS 失败/为空时,根据 `anti_bot_mode` 使用 stealth 或 headful Playwright。 - **Profile 覆盖**:`EN_NEWS_PROFILE` 环境变量可切换 `8g_headful` / `2g_headless` 等抓取配置。 ### 5.2 三层去重 | 层级 | 维度 | 说明 | |------|------|------| | L1 | URL | 相同 URL 直接判重 | | L2 | 内容 | 正文 hash 相同判重 | | L3 | SimHash | 汉明距离 ≤ 阈值判为模糊重复 | 指纹库为 `data/dedup/fingerprints.sqlite3`,SimHash 窗口默认 30 天。 ### 5.3 LLM 场景配置 `configs/system.yaml` 中通过 `llm_scenes` 为不同任务配置独立模型: | 场景 | 用途 | 当前默认 | |------|------|----------| | `translation` | M4 翻译+事件抽取 | deepseek-v4-flash,temperature=0.1,max_tokens=8192 | | `daily_report` | M7 日报 AI 摘要 | deepseek-v4-flash,temperature=0.3,max_tokens=1500 | 未在场景中声明的字段自动回退到 `llm` 默认段。Embedding 是单一场景,不能按任务拆分(查询/入库向量必须同模型)。 ### 5.4 增量与幂等 | 阶段 | 增量机制 | |------|----------| | M2 | `processed/{source}/{date}/{url_hash}.json` 存在则跳过 | | M3 | 指纹库判重;重复篇不再写入,只合并来源 | | M4 | `events/{date}/{url_hash}.json` 存在则跳过 | | M5 | `embeddings/{date}/{url_hash}.json` 存在则跳过 | | M6 | Qdrant point ID = url_hash 的 UUID,重复 upsert 覆盖 | | M9 | MySQL 唯一键 `(report_date, report_type='intl', file_name='')` 覆盖 | Shell 层还提供步骤级 `--resume`:状态写在 `data/run_state/{YYYYMMDD}.state`。 ### 5.5 日报生成模式 - 时间窗口:过去 25 小时(跨天自动聚合)。 - 高重要度事件:`importance >= 4`,不足 3 条时逐级降阈。 - AI 摘要:事件超过 10 条自动分批生成,再合并;LLM 失败走规则兜底,不阻断日报入库。 - 日报不再生成 HTML/上传,而是结构化写入 MySQL。 ## 6. 部署拓扑(2026-07 起) ```text 海外服务器(已停用) ↓ 历史:抓取后同步回国内 国内服务器 Pi(当前唯一全链路节点) /home/pi/intlnews ├── Xvfb :99(headful 浏览器虚拟显示) ├── Privoxy → SOCKS5(HTTP 代理访问海外) ├── crontab: 06:00 / 12:00 / 18:00 / 22:00 执行 domestic_full.sh ├── 本地 Qdrant 文件模式: data/qdrant_storage └── MySQL: 通过内网隧道访问 news 项目共用 myquant 库 ``` ## 7. 关键目录 ```text app/ CLI 入口 crawler/ M1 抓取 extractor/ M2 正文提取 dedup/ M3 去重 llm/ M4 翻译/事件抽取 embedding/ M5 向量化 vectorstore/ M6 Qdrant scheduler/ M7/M9 编排与日报 report_db/ M9 MySQL mcp_server/ M8 MCP prompts/ LLM Prompt 模板 configs/ 系统/源/Profile 配置 scripts/ Bash 自动化脚本 tests/ 单元测试 docs/ 本文档目录 ```