Files
simon 3b44f64f66 refactor: 清理历史 AI Agent 文档残留 + 重构 docs/ + Pipeline 健壮性修复
- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
        统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明

Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
2026-08-22 20:47:53 +08:00

172 lines
7.7 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 项目架构说明
## 1. 定位
`en-news` 是一个从英文财经新闻源采集数据,经过正文提取、去重、LLM 翻译/事件抽取、向量化、向量检索,最终生成中文日报的私有化工作流。
设计目标:
- **私有化**:不依赖海外服务器,抓取与计算均可在国内 Raspberry Pi 上独立运行。
- **增量友好**:每个阶段都尽量跳过已处理文件,断点续跑成本低。
- **幂等**:重复运行不会产生重复数据或重复日报。
- **可运维**:终端输出阶段/模型/耗时,日志统一写入 `logs/`。
## 2. 总体架构
```text
┌─────────────────────────────────────────────────────────────┐
│ en-news 平台 │
│ │
│ configs/ · .env · prompts/ · data/ · logs/ │
│ │
│ app/cli.py ── Typer CLI 入口 │
│ ├── crawl M1 抓取 │
│ ├── extract M2 正文提取 │
│ ├── dedup M3 三层去重 │
│ ├── translate M4 翻译 + 事件抽取 │
│ ├── embed M5 向量生成 │
│ ├── index M6 Qdrant 入库 │
│ ├── search M6 语义搜索 │
│ ├── report M9 日报结构化入库 │
│ ├── pipeline M2→M6(+日报)一键管道 │
│ └── mcp-server M8 MCP 服务 │
│ │
│ scheduler/ ── 管道编排 / 日报生成 │
│ mcp_server/ ── MCP 工具层(研究 Agent 接入) │
└─────────────────────────────────────────────────────────────┘
```
## 3. 模块职责
| 模块 | 对应阶段 | 核心职责 |
|------|----------|----------|
| `crawler/` | M1 | 从 RSS/Atom/网页抓取新闻,RSS 优先、Stealth/Headful Web 回退,写入 `data/raw/` |
| `extractor/` | M2 | 使用 trafilatura(含 Markdown 回退)提取正文,输出 `data/processed/` |
| `dedup/` | M3 | URL、内容、SimHash 三层去重,SQLite 指纹库,跨源来源合并,输出 `data/deduped/` |
| `llm/` | M4 | 调用 DeepSeek/Qwen 兼容接口,单次调用完成全文英译中+投资事件抽取,输出 `data/events/` |
| `embedding/` | M5 | 调用 DashScope text-embedding 批量向量化,输出 `data/embeddings/` |
| `vectorstore/` | M6 | 管理 Qdrant collection,幂等 upsert、语义搜索、过滤查询 |
| `scheduler/` | M7/M9 | 串联 M2→M6,生成每日 AI 摘要日报并写入 MySQL |
| `mcp_server/` | M8 | 暴露 MCP 工具供 Cherry Studio/Claude Code 等调用 |
| `report_db/` | M9 | MySQL `news_report` / `news_event` 建表与读写 |
| `app/cli.py` | 入口 | Typer CLI,各模块的同步调用入口 |
| `scripts/` | 运维 | Bash 全流程、断点续跑、抓取 Profile、日志清理 |
## 4. 数据流
```text
M1 抓取
data/raw/{source_id}/{YYYYMMDD}/
├── index.jsonl # 抓取索引(每条含 url_hash、标题、URL)
├── {url_hash}.html # 原始 HTML
└── {url_hash}.md # Crawl4AI 可选的 Markdown
M2 正文提取
data/processed/{source_id}/{YYYYMMDD}/{url_hash}.json
# ProcessedArticle:标题、正文、词数、提取器、时间等
M3 三层去重
data/deduped/{YYYYMMDD}/uniques/{url_hash}.json
# 跨源重复时在已保留唯一篇的 source_ids 中追加来源
M4 翻译 + 事件抽取
data/events/{YYYYMMDD}/{url_hash}.json
# EnTranslatedArticle:中英文标题/正文、事件列表、source_ids
M5 向量生成
data/embeddings/{YYYYMMDD}/{url_hash}.json
# 向量文件(1024 维 text-embedding-v3)+ 每日期 index.json
M6 Qdrant 入库
vectorstore client → collection: en_finance_news
# payload 含标题、双语标题、事件、来源、时间、正文预览
M7/M9 日报
scheduler/reporter.py → MySQL:
news_report(主表:日期、type=intl、AI 摘要、stats)
news_event(明细:Top 事件、来源、情绪、重要度、URL)
```
## 5. 关键技术点
### 5.1 抓取策略
- **RSS 优先**:大部分源先用 RSS/Atom/Google News RSS,减少对浏览器的依赖。
- **Web 回退**:RSS 失败/为空时,根据 `anti_bot_mode` 使用 stealth 或 headful Playwright。
- **Profile 覆盖**:`EN_NEWS_PROFILE` 环境变量可切换 `8g_headful` / `2g_headless` 等抓取配置。
### 5.2 三层去重
| 层级 | 维度 | 说明 |
|------|------|------|
| L1 | URL | 相同 URL 直接判重 |
| L2 | 内容 | 正文 hash 相同判重 |
| L3 | SimHash | 汉明距离 ≤ 阈值判为模糊重复 |
指纹库为 `data/dedup/fingerprints.sqlite3`,SimHash 窗口默认 30 天。
### 5.3 LLM 场景配置
`configs/system.yaml` 中通过 `llm_scenes` 为不同任务配置独立模型:
| 场景 | 用途 | 当前默认 |
|------|------|----------|
| `translation` | M4 翻译+事件抽取 | deepseek-v4-flash,temperature=0.1,max_tokens=8192 |
| `daily_report` | M7 日报 AI 摘要 | deepseek-v4-flash,temperature=0.3,max_tokens=1500 |
未在场景中声明的字段自动回退到 `llm` 默认段。Embedding 是单一场景,不能按任务拆分(查询/入库向量必须同模型)。
### 5.4 增量与幂等
| 阶段 | 增量机制 |
|------|----------|
| M2 | `processed/{source}/{date}/{url_hash}.json` 存在则跳过 |
| M3 | 指纹库判重;重复篇不再写入,只合并来源 |
| M4 | `events/{date}/{url_hash}.json` 存在则跳过 |
| M5 | `embeddings/{date}/{url_hash}.json` 存在则跳过 |
| M6 | Qdrant point ID = url_hash 的 UUID,重复 upsert 覆盖 |
| M9 | MySQL 唯一键 `(report_date, report_type='intl', file_name='')` 覆盖 |
Shell 层还提供步骤级 `--resume`:状态写在 `data/run_state/{YYYYMMDD}.state`。
### 5.5 日报生成模式
- 时间窗口:过去 25 小时(跨天自动聚合)。
- 高重要度事件:`importance >= 4`,不足 3 条时逐级降阈。
- AI 摘要:事件超过 10 条自动分批生成,再合并;LLM 失败走规则兜底,不阻断日报入库。
- 日报不再生成 HTML/上传,而是结构化写入 MySQL。
## 6. 部署拓扑(2026-07 起)
```text
海外服务器(已停用)
↓ 历史:抓取后同步回国内
国内服务器 Pi(当前唯一全链路节点)
/home/pi/intlnews
├── Xvfb :99(headful 浏览器虚拟显示)
├── Privoxy → SOCKS5(HTTP 代理访问海外)
├── crontab: 06:00 / 12:00 / 18:00 / 22:00 执行 domestic_full.sh
├── 本地 Qdrant 文件模式: data/qdrant_storage
└── MySQL: 通过内网隧道访问 news 项目共用 myquant 库
```
## 7. 关键目录
```text
app/ CLI 入口
crawler/ M1 抓取
extractor/ M2 正文提取
dedup/ M3 去重
llm/ M4 翻译/事件抽取
embedding/ M5 向量化
vectorstore/ M6 Qdrant
scheduler/ M7/M9 编排与日报
report_db/ M9 MySQL
mcp_server/ M8 MCP
prompts/ LLM Prompt 模板
configs/ 系统/源/Profile 配置
scripts/ Bash 自动化脚本
tests/ 单元测试
docs/ 本文档目录
```