refactor: 清理历史 AI Agent 文档残留 + 重构 docs/ + Pipeline 健壮性修复

- docs: 删除 CLAUDE.md / continuation.md / english-news-plan.md 及旧版 intlnews_usage.*,
        统一迁移到 docs/{README,architecture,quickstart,usage,pipeline,configuration,deployment,development,faq}.md
- README: 精简为仓库入口,指向 docs/
- configs/sources.yaml: 更新注释指向新文档
- .env.example: 修正 DashScope Embedding 端点说明

Pipeline 修复:
- dedup/llm/embedding/vectorstore/reporter: 过滤 M2 no_content / 空正文,避免污染下游与 Qdrant
- dedup/pipeline: 改为先写唯一文件再写指纹,避免崩溃导致文章永久丢失
- crawler/orchestrator: sources_crawled 改为“尝试数”,成功数 = crawled - failed
- crawler/storage: write_index_jsonl 从文章路径推断日期,修复跨天/测试路径问题
- scheduler/pipeline: STEP_TIMEOUTS 实际生效(SIGALRM)
- scheduler/reporter: emb_count 排除 index.json;日报跳过无原文事件
- vectorstore/pipeline: payload 增加 source_ids;--recreate --all 时空日期也重建 collection
- app/cli: extract/dedup/translate/embed/index/pipeline 支持 --date;embed/index 支持 --all;crawl 全源失败返回非零
- scripts: domestic_full/crawl_8g/crawl_2g/pipeline 安全加载 .env;M1 全失败不标记且最终退出码=1
This commit is contained in:
2026-08-22 20:47:53 +08:00
parent 951276a313
commit 3b44f64f66
30 changed files with 1496 additions and 2804 deletions
+27
View File
@@ -0,0 +1,27 @@
# 文档中心
本项目 **English Financial News**(`en-news`)是一套面向国际财经新闻的私有化 Deep Research 平台:抓取、去重、翻译、事件抽取、向量化、语义检索、日报生成全部在本地/Pi 上独立运行。
> 本目录是项目文档的唯一入口。历史 AI Agent 文档(`CLAUDE.md`、`continuation.md`、`english-news-plan.md`、旧版 `docs/intlnews_usage.*`)已清理合并到以下文档中。
## 文档导航
| 文档 | 内容 |
|------|------|
| [架构说明](architecture.md) | 总体架构、模块职责、数据流、部署拓扑 |
| [快速开始](quickstart.md) | 环境准备、安装、配置、首次运行 |
| [使用手册](usage.md) | CLI 命令、Shell 脚本、MCP 服务、日报查看 |
| [流水线详解](pipeline.md) | M1→M9 各步骤输入/输出、增量逻辑、幂等机制 |
| [配置说明](configuration.md) | `system.yaml`、`sources.yaml`、Profile、`.env` |
| [部署与运维](deployment.md) | Pi 服务器部署、Xvfb/代理、crontab、日志清理 |
| [开发指南](development.md) | 源码结构、测试、已知问题、新增新闻源 |
| [FAQ](faq.md) | 常见问题与排查 |
## 顶层入口
- 项目根目录的 [README.md](../README.md) 仅保留简短的仓库入口和文档链接。
- 当前推荐的执行入口:
- 全流程:`bash scripts/domestic_full.sh`
- 仅 M2→M6+日报:`bash scripts/pipeline.sh`
- CLI 单步:`uv run en-news <command>`
- MCP:`uv run en-news mcp-server`
+171
View File
@@ -0,0 +1,171 @@
# 项目架构说明
## 1. 定位
`en-news` 是一个从英文财经新闻源采集数据,经过正文提取、去重、LLM 翻译/事件抽取、向量化、向量检索,最终生成中文日报的私有化工作流。
设计目标:
- **私有化**:不依赖海外服务器,抓取与计算均可在国内 Raspberry Pi 上独立运行。
- **增量友好**:每个阶段都尽量跳过已处理文件,断点续跑成本低。
- **幂等**:重复运行不会产生重复数据或重复日报。
- **可运维**:终端输出阶段/模型/耗时,日志统一写入 `logs/`。
## 2. 总体架构
```text
┌─────────────────────────────────────────────────────────────┐
│ en-news 平台 │
│ │
│ configs/ · .env · prompts/ · data/ · logs/ │
│ │
│ app/cli.py ── Typer CLI 入口 │
│ ├── crawl M1 抓取 │
│ ├── extract M2 正文提取 │
│ ├── dedup M3 三层去重 │
│ ├── translate M4 翻译 + 事件抽取 │
│ ├── embed M5 向量生成 │
│ ├── index M6 Qdrant 入库 │
│ ├── search M6 语义搜索 │
│ ├── report M9 日报结构化入库 │
│ ├── pipeline M2→M6(+日报)一键管道 │
│ └── mcp-server M8 MCP 服务 │
│ │
│ scheduler/ ── 管道编排 / 日报生成 │
│ mcp_server/ ── MCP 工具层(研究 Agent 接入) │
└─────────────────────────────────────────────────────────────┘
```
## 3. 模块职责
| 模块 | 对应阶段 | 核心职责 |
|------|----------|----------|
| `crawler/` | M1 | 从 RSS/Atom/网页抓取新闻,RSS 优先、Stealth/Headful Web 回退,写入 `data/raw/` |
| `extractor/` | M2 | 使用 trafilatura(含 Markdown 回退)提取正文,输出 `data/processed/` |
| `dedup/` | M3 | URL、内容、SimHash 三层去重,SQLite 指纹库,跨源来源合并,输出 `data/deduped/` |
| `llm/` | M4 | 调用 DeepSeek/Qwen 兼容接口,单次调用完成全文英译中+投资事件抽取,输出 `data/events/` |
| `embedding/` | M5 | 调用 DashScope text-embedding 批量向量化,输出 `data/embeddings/` |
| `vectorstore/` | M6 | 管理 Qdrant collection,幂等 upsert、语义搜索、过滤查询 |
| `scheduler/` | M7/M9 | 串联 M2→M6,生成每日 AI 摘要日报并写入 MySQL |
| `mcp_server/` | M8 | 暴露 MCP 工具供 Cherry Studio/Claude Code 等调用 |
| `report_db/` | M9 | MySQL `news_report` / `news_event` 建表与读写 |
| `app/cli.py` | 入口 | Typer CLI,各模块的同步调用入口 |
| `scripts/` | 运维 | Bash 全流程、断点续跑、抓取 Profile、日志清理 |
## 4. 数据流
```text
M1 抓取
data/raw/{source_id}/{YYYYMMDD}/
├── index.jsonl # 抓取索引(每条含 url_hash、标题、URL)
├── {url_hash}.html # 原始 HTML
└── {url_hash}.md # Crawl4AI 可选的 Markdown
M2 正文提取
data/processed/{source_id}/{YYYYMMDD}/{url_hash}.json
# ProcessedArticle:标题、正文、词数、提取器、时间等
M3 三层去重
data/deduped/{YYYYMMDD}/uniques/{url_hash}.json
# 跨源重复时在已保留唯一篇的 source_ids 中追加来源
M4 翻译 + 事件抽取
data/events/{YYYYMMDD}/{url_hash}.json
# EnTranslatedArticle:中英文标题/正文、事件列表、source_ids
M5 向量生成
data/embeddings/{YYYYMMDD}/{url_hash}.json
# 向量文件(1024 维 text-embedding-v3)+ 每日期 index.json
M6 Qdrant 入库
vectorstore client → collection: en_finance_news
# payload 含标题、双语标题、事件、来源、时间、正文预览
M7/M9 日报
scheduler/reporter.py → MySQL:
news_report(主表:日期、type=intl、AI 摘要、stats)
news_event(明细:Top 事件、来源、情绪、重要度、URL)
```
## 5. 关键技术点
### 5.1 抓取策略
- **RSS 优先**:大部分源先用 RSS/Atom/Google News RSS,减少对浏览器的依赖。
- **Web 回退**:RSS 失败/为空时,根据 `anti_bot_mode` 使用 stealth 或 headful Playwright。
- **Profile 覆盖**:`EN_NEWS_PROFILE` 环境变量可切换 `8g_headful` / `2g_headless` 等抓取配置。
### 5.2 三层去重
| 层级 | 维度 | 说明 |
|------|------|------|
| L1 | URL | 相同 URL 直接判重 |
| L2 | 内容 | 正文 hash 相同判重 |
| L3 | SimHash | 汉明距离 ≤ 阈值判为模糊重复 |
指纹库为 `data/dedup/fingerprints.sqlite3`,SimHash 窗口默认 30 天。
### 5.3 LLM 场景配置
`configs/system.yaml` 中通过 `llm_scenes` 为不同任务配置独立模型:
| 场景 | 用途 | 当前默认 |
|------|------|----------|
| `translation` | M4 翻译+事件抽取 | deepseek-v4-flash,temperature=0.1,max_tokens=8192 |
| `daily_report` | M7 日报 AI 摘要 | deepseek-v4-flash,temperature=0.3,max_tokens=1500 |
未在场景中声明的字段自动回退到 `llm` 默认段。Embedding 是单一场景,不能按任务拆分(查询/入库向量必须同模型)。
### 5.4 增量与幂等
| 阶段 | 增量机制 |
|------|----------|
| M2 | `processed/{source}/{date}/{url_hash}.json` 存在则跳过 |
| M3 | 指纹库判重;重复篇不再写入,只合并来源 |
| M4 | `events/{date}/{url_hash}.json` 存在则跳过 |
| M5 | `embeddings/{date}/{url_hash}.json` 存在则跳过 |
| M6 | Qdrant point ID = url_hash 的 UUID,重复 upsert 覆盖 |
| M9 | MySQL 唯一键 `(report_date, report_type='intl', file_name='')` 覆盖 |
Shell 层还提供步骤级 `--resume`:状态写在 `data/run_state/{YYYYMMDD}.state`。
### 5.5 日报生成模式
- 时间窗口:过去 25 小时(跨天自动聚合)。
- 高重要度事件:`importance >= 4`,不足 3 条时逐级降阈。
- AI 摘要:事件超过 10 条自动分批生成,再合并;LLM 失败走规则兜底,不阻断日报入库。
- 日报不再生成 HTML/上传,而是结构化写入 MySQL。
## 6. 部署拓扑(2026-07 起)
```text
海外服务器(已停用)
↓ 历史:抓取后同步回国内
国内服务器 Pi(当前唯一全链路节点)
/home/pi/intlnews
├── Xvfb :99(headful 浏览器虚拟显示)
├── Privoxy → SOCKS5(HTTP 代理访问海外)
├── crontab: 06:00 / 12:00 / 18:00 / 22:00 执行 domestic_full.sh
├── 本地 Qdrant 文件模式: data/qdrant_storage
└── MySQL: 通过内网隧道访问 news 项目共用 myquant 库
```
## 7. 关键目录
```text
app/ CLI 入口
crawler/ M1 抓取
extractor/ M2 正文提取
dedup/ M3 去重
llm/ M4 翻译/事件抽取
embedding/ M5 向量化
vectorstore/ M6 Qdrant
scheduler/ M7/M9 编排与日报
report_db/ M9 MySQL
mcp_server/ M8 MCP
prompts/ LLM Prompt 模板
configs/ 系统/源/Profile 配置
scripts/ Bash 自动化脚本
tests/ 单元测试
docs/ 本文档目录
```
+238
View File
@@ -0,0 +1,238 @@
# 配置说明
## 1. 配置文件总览
| 文件 | 用途 |
|------|------|
| `configs/system.yaml` | 系统功能配置:代理、抓取、提取、去重、LLM、Embedding、Qdrant、日报、调度、日志 |
| `configs/sources.yaml` | 新闻源列表与抓取规则 |
| `configs/profiles/8g_headful.yaml` | Pi 8G 有头浏览器 + HTTP 代理配置 |
| `configs/profiles/2g_headless.yaml` | 2G 轻量 headless stealth 配置 |
| `.env` | 敏感配置:API Key、数据库密码、远程 Qdrant 地址 |
| `.env.example` | 环境变量模板 |
| `prompts/translation_and_extraction.md` | M4 翻译+事件抽取 Prompt 模板 |
## 2. `.env` 环境变量
```env
# DeepSeek(M4 / M7 默认 LLM)
DEEPSEEK_API_KEY=sk-your-deepseek-key
DEEPSEEK_BASE_URL=https://api.deepseek.com
# Qwen(可选,切换 LLM provider 时使用)
QWEN_API_KEY=sk-your-qwen-key
QWEN_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
# DashScope(M5 Embedding)
# 使用 OpenAI 兼容端点,默认复用 QWEN_BASE_URL;如需独立端点可设置 DASHSCOPE_EMBEDDING_BASE_URL
DASHSCOPE_API_KEY=sk-your-dashscope-key
# DASHSCOPE_EMBEDDING_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
# Qdrant(可选;默认本地文件模式)
QDRANT_URL=http://localhost:6333
QDRANT_API_KEY=
# MySQL(M9 日报入库)
NEWS_DB_HOST=192.168.1.10
NEWS_DB_PORT=13306
NEWS_DB_USER=myquant
NEWS_DB_PASSWORD=
NEWS_DB_NAME=myquant
```
## 3. `configs/system.yaml`
### 3.1 servers
```yaml
servers:
overseas_host: "ecs-user@8.217.19.253"
overseas_path: "/opt/intlgrab"
domestic_host: "pi@192.168.1.160"
domestic_path: "/home/pi/intlnews"
```
海外服务器已停用,字段保留为兼容参考。
### 3.2 proxy
```yaml
proxy:
enabled: false
url: "http://127.0.0.1:3128"
bypass_domains: []
```
- 默认关闭。
- `8g_headful` Profile 会开启并通过 `http://127.0.0.1:3128`(Privoxy → SOCKS5)访问海外。
### 3.3 crawler
| 字段 | 默认 | 说明 |
|------|------|------|
| `max_memory_mb` | 1800 | 内存上限,超限触发 GC/保护逻辑 |
| `source_timeout_sec` | 7200 | 单源超时 |
| `article_delay_sec` | 3.0 | 文章间冷却 |
| `page_timeout_sec` | 45 | 单页加载超时 |
| `viewport_width/height` | 1024/768 | 浏览器视口 |
| `headful` | false | true=有头浏览器 |
| `xvfb_display` | ":99" | Xvfb 虚拟显示器 |
### 3.4 extractor / dedup
```yaml
extractor:
min_content_words: 50
trafilatura_fallback: true
dedup:
hamming_distance_threshold: 3
simhash_window_days: 30
min_content_length: 100
```
### 3.5 llm 与 llm_scenes
```yaml
llm:
provider: "deepseek"
deepseek_model: "deepseek-v4-flash"
qwen_model: "qwen-plus"
timeout_sec: 60
max_attempts: 3
max_tokens: 8192
temperature: 0.1
concurrency: 3
llm_scenes:
translation:
provider: "deepseek"
model: "deepseek-v4-flash"
temperature: 0.1
max_tokens: 8192
daily_report:
provider: "deepseek"
model: "deepseek-v4-flash"
temperature: 0.3
max_tokens: 1500
```
- `translation`:M4 使用。
- `daily_report`:M7 日报摘要使用。
- 场景未声明的字段会回退 `llm`。
### 3.6 embedding
```yaml
embedding:
provider: "dashscope"
dashscope_model: "text-embedding-v3"
dimension: 1024
batch_size: 10
max_attempts: 3
timeout_sec: 30
```
注意:Embedding 不支持按场景拆分配置,因为入库向量与查询向量必须同模型。
### 3.7 qdrant / report / schedule / logging
```yaml
qdrant:
collection: "en_finance_news"
report:
max_events: 20
summary_max_chars: 500
importance_threshold: 4
upload_host: "simon@doorcome.cn"
upload_path: "/var/www/html/echart/research"
schedule:
day_cutoff_hour: 6
times: ["06:00", "12:00", "18:00", "22:00"]
logging:
level: "INFO"
dir: "logs"
```
`report.upload_host/upload_path` 是 M9 之前 HTML 日报上传的旧配置,当前仅作保留。
## 4. `configs/sources.yaml`
### 4.1 顶层 settings
```yaml
settings:
concurrency: 5
request_delay_sec: 2
user_agent: "..."
```
### 4.2 单源字段
| 字段 | 说明 |
|------|------|
| `id` | 唯一标识,如 `reuters` |
| `name` | 展示名,如 `Reuters` |
| `enabled` | 是否启用 |
| `homepage` | 主页,用于回退/域名校验 |
| `article_url_pattern` | URL 正则,过滤非文章链接 |
| `js_render` | 是否需要 JS 渲染 |
| `max_articles_per_run` | 单次最多抓取篇数 |
| `rss_url` | RSS/Atom/Google News RSS 地址 |
| `anti_bot_mode` | `stealth` / `headful` 等 Web 回退模式 |
### 4.3 当前源列表
| ID | 名称 | 抓取方式 | 状态 |
|----|------|----------|------|
| `reuters` | Reuters | Google News RSS(可能只有摘要) | ⚠️ DataDome |
| `cnbc` | CNBC | CNBC RSS | ✅ |
| `marketwatch` | MarketWatch | RSS + stealth 回退 | ✅ |
| `ft` | Financial Times | FT RSS(摘要) | ⚠️ 付费墙 |
| `yahoo_finance` | Yahoo Finance | Yahoo RSS | ✅ |
| `investing` | Investing.com | Google News RSS + stealth | ✅(摘要) |
| `seekingalpha` | Seeking Alpha | RSS + stealth | ✅ |
| `barrons` | Barron's | Dow Jones RSS + stealth | ✅ |
| `wsj` | WSJ | Dow Jones RSS + stealth | ✅ |
| `economist` | The Economist | RSS + stealth | ✅ |
| `investinglive` | InvestingLive | RSS(全文) | ✅ |
| `zerohedge` | ZeroHedge | FeedBurner RSS(全文) | ✅ |
> 历史 `forexlive` 已迁移为 `investinglive`。
## 5. Profiles
### 5.1 `8g_headful.yaml`
适合国内 Pi(内存充裕):
- 开启 HTTP 代理
- `headful: true`
- 更大视口/超时
- 通过 Xvfb `:99` 运行
### 5.2 `2g_headless.yaml`
适合低配/海外轻量:
- headless stealth
- 内存上限 1800MB
- 默认不开启代理
通过环境变量切换:
```bash
EN_NEWS_PROFILE=8g_headful uv run en-news crawl
# 或直接使用对应脚本
bash scripts/domestic_crawl_8g.sh
```
## 6. Prompt 模板
`prompts/translation_and_extraction.md` 包含:
- `## System Prompt`:LLM System 指令
- `## User Input`:带 `{title}`、`{source_name}`、`{publish_time}`、`{content}` 占位符
如需调整事件类型、输出格式或翻译风格,应修改此文件并同步测试。
+145
View File
@@ -0,0 +1,145 @@
# 部署与运维
## 1. 目标环境
当前生产环境为国内单节点(Raspberry Pi 4, 8GB),路径 `/home/pi/intlnews`。海外服务器已停用。
```text
Pi /home/pi/intlnews
├── Python 3.11 + uv + .venv
├── Playwright(Chromium)
├── Xvfb :99
├── Privoxy 127.0.0.1:3128 → ss-local SOCKS5 1088
├── 本地 Qdrant 文件模式 data/qdrant_storage
└── crontab 定时 full pipeline
```
## 2. 系统依赖
```bash
sudo apt update
sudo apt install -y python3.11 python3.11-venv uv xvfb \
shadowsocks-libev privoxy chromium # chromium 可根据 Playwright 安装方式调整
```
> 本项目使用 `uv` 管理 Python 环境,不依赖系统 pip 安装依赖。Playwright 浏览器建议按 `crawl4ai` 文档安装。
## 3. 首次部署步骤
```bash
# 1. 拉取/同步代码到 /home/pi/intlnews
cd /home/pi/intlnews
# 2. 安装 Python 依赖
uv sync
# 3. 准备 .env
cp .env.example .env
# 编辑 .env:DeepSeek / DashScope / MySQL 等
# 4. 启动 Xvfb(headful 抓取需要)
Xvfb :99 -screen 0 1280x1024x24 -ac +extension RANDR &
# 5. 启动代理(如果国内访问海外需要)
sudo systemctl enable --now shadowsocks-libev-local
sudo systemctl enable --now privoxy
# 验证代理
curl -x http://127.0.0.1:3128 -I https://www.google.com
```
## 4. 代理配置参考
- Shadowsocks 配置文件:`/etc/shadowsocks-libev/config.json`
- Privoxy 配置追加:
```
forward-socks5t 127.0.0.1:1088 .
listen-address 127.0.0.1:3128
```
- `domestic_crawl_8g.sh` 会自动 `export HTTP_PROXY/HTTPS_PROXY` 到 `127.0.0.1:3128`,并使用 `8g_headful` Profile。
## 5. Xvfb 管理
```bash
# 查看是否运行
pgrep -x Xvfb
# 手动启动
Xvfb :99 -screen 0 1280x1024x24 -ac +extension RANDR &
# 开机自启(示例 systemd)
# 也可由 domestic_crawl_8g.sh 自动检测并启动
```
## 6. 定时任务(crontab)
推荐 crontab:
```cron
0 6 * * * cd /home/pi/intlnews && bash scripts/domestic_full.sh >> logs/cron_06.log 2>&1
0 12 * * * cd /home/pi/intlnews && bash scripts/domestic_full.sh >> logs/cron_12.log 2>&1
0 18 * * * cd /home/pi/intlnews && bash scripts/domestic_full.sh >> logs/cron_18.log 2>&1
0 22 * * * cd /home/pi/intlnews && bash scripts/domestic_full.sh >> logs/cron_22.log 2>&1
```
如果希望中断恢复语义,可以在每次调度前追加 `--resume`?注意:`--resume` 通常用于手动恢复;日常定时全量不带 `--resume` 也能靠文件级增量避免重复处理。按需选择。
## 7. MySQL / 日报依赖
M9 日报写入共用 MySQL `myquant` 库,需要以下网络配置:
- Pi 通过内网访问 `192.168.1.10:13306`(该端口是到远程 MySQL 的 autossh 隧道)
- 用户名/库名通常为 `myquant` / `myquant`
- `.env` 中必须配置 `NEWS_DB_PASSWORD`
验证:
```bash
uv run en-news report
```
成功会输出 `report_id=...`。
## 8. 日志与清理
- 日志目录:`logs/`
- `domestic_full_*.log`
- `pipeline_*.log`
- 保留策略:默认保留 14 天,执行 `scripts/cleanup_logs.sh`。
- 可加入 crontab:
```cron
30 3 * * * cd /home/pi/intlnews && bash scripts/cleanup_logs.sh >> /dev/null 2>&1
```
## 9. 升级/同步代码
```bash
cd /home/pi/intlnews
# 拉取新代码
git pull # 如果使用 git
# 同步依赖
uv sync
# 执行一次冒烟
uv run en-news --help
# 查看关键日志
tail -100 logs/pipeline_*.log | tail -100
```
## 10. 故障排查指引
| 现象 | 可能原因 | 处理 |
|------|----------|------|
| 抓取全部失败 | 代理未启动 / 网络不通 | `curl -x http://127.0.0.1:3128 ...` |
| headful 起不来 | Xvfb 未运行 / DISPLAY 不对 | 启动 Xvfb,`DISPLAY=:99` |
| M4 翻译失败 | DeepSeek Key 未配置 / 余额不足 | 检查 `.env`,单独跑 `uv run en-news translate` |
| M5 失败 | DashScope Key 未配置 | 检查 `.env` |
| report 失败 | `NEWS_DB_PASSWORD` 缺失 / 隧道不通 | 检查 `.env` 和 MySQL 端口 |
| Qdrant 入库失败 | 本地文件损坏 / collection 异常 | 备份后重建 `data/qdrant_storage`,`uv run en-news index --recreate` |
## 11. 容量建议
- `data/` 会随时间增长,建议定期归档旧数据。
- 日志 `logs/` 用 `cleanup_logs.sh` 清理。
- Qdrant 本地文件模式对 Pi 友好,但大规模检索建议迁移到独立服务器/远程 Qdrant。
+101
View File
@@ -0,0 +1,101 @@
# 开发指南
## 1. 代码结构
```text
app/cli.py Typer CLI 入口
crawler/ M1 抓取
config.py 系统配置 + Profile 深度合并
orchestrator.py 串行抓取编排、RSS 优先 Web 回退
rss_crawler.py RSS/Atom/Google News RSS 抓取
crawler.py Crawl4AI Web 抓取
loader.py 读取 sources.yaml
storage.py index.jsonl 存储
extractor/ M2 正文提取
dedup/ M3 三层去重
llm/ M4 LLM 翻译/事件抽取
embedding/ M5 向量化
vectorstore/ M6 Qdrant
scheduler/ M7/M9 管道编排与日报
report_db/ M9 MySQL
mcp_server/ M8 MCP
prompts/ LLM Prompt 模板
scripts/ Bash 自动化与运维
tests/ 单元测试
docs/ 项目文档
```
## 2. 常用开发命令
```bash
# 安装开发依赖
uv sync --extra dev
# 运行全部测试
uv run pytest
# 运行单个测试文件
uv run pytest tests/test_llm.py
# 代码检查
uv run ruff check .
# 自动修复
uv run ruff check . --fix
# 运行 CLI
uv run en-news --help
```
## 3. 测试现状
- 当前测试覆盖:crawler、extractor、dedup、llm、embedding、vectorstore、scheduler、report_db、mcp_server。
- 已知问题:
- `tests/test_crawler.py::test_write_and_load_index_jsonl`
- `tests/test_crawler.py::test_write_index_jsonl_dedup`
- 原因是 `crawler/storage.py:load_index()` 硬编码 `data/raw/...`,未使用测试临时目录。
- `scheduler/reporter.py` 与 `report_db/schema.py` 因内含长模板/DDL,已在 `pyproject.toml` 中忽略 E501。
## 4. 新增新闻源
1. 在 `configs/sources.yaml` 的 `sources:` 列表新增条目:
```yaml
- id: "example"
name: "Example News"
enabled: true
homepage: "https://www.example.com/"
article_url_pattern: "/news/"
js_render: false
max_articles_per_run: 20
rss_url: "https://www.example.com/rss"
anti_bot_mode: "stealth"
```
2. 本地测试:
```bash
uv run en-news crawl --source example
uv run en-news extract --source example
```
3. 确认 `data/raw/example/...` 和 `data/processed/example/...` 正常。
4. 如果使用 Google News RSS,需确认 `rss_crawler.py` 能正确清理标题和识别跳转链接。
## 5. LLM Prompt 开发
M4 Prompt 位于 `prompts/translation_and_extraction.md`。
- 不要随意改变 JSON 输出结构,否则 `llm/models.py::LLMTranslationOutput` 可能校验失败。
- 事件类型定义在 `llm/models.py::INTERNATIONAL_EVENT_TYPES`,Prompt 中应保持一致。
- 修改后建议跑 `uv run pytest tests/test_llm.py`。
## 6. 数据模型核心字段
- `ProcessedArticle`(`extractor/models.py`):包含 `source_id`、`source_ids`、标题、URL、正文、词数等。
- `EnTranslatedArticle`(`llm/models.py`):包含中英文标题/正文、`events`、`source_ids`。
- `EventExtraction`(`llm/models.py`):事件类型、代码、情绪、重要度、摘要。
- `ReportData` / `EventRow`(`report_db/models.py`):日报入库结构。
## 7. 文档维护规范
- 所有面向使用/部署/开发的 Markdown 文档统一放在 `docs/`。
- 根目录 `README.md` 仅保留仓库入口和文档链接。
- 新增功能或修改流程时,同步更新对应 `docs/*.md`。
- 不再在仓库根目录维护 `CLAUDE.md` / `continuation.md` / `english-news-plan.md` 这类 AI Agent 会话残留。
+104
View File
@@ -0,0 +1,104 @@
# FAQ 常见问题
## 1. 为什么有些源只拿到摘要?
部分网站存在 Cloudflare / DataDome / 付费墙等反爬限制:
- RSS 源只返回标题+摘要(如 Reuters、Investing.com、FT)
- Web 全文抓取可能被拦截
- 当前策略是 RSS 优先 + stealth/headful Web 回退,但无法保证 100% 全文
解决方法:
- 使用质量更好的代理/住宅 IP
- 订阅相应媒体付费内容
- 对 Pipeline 而言,摘要也能进入翻译/抽取,只是信息量较低
## 2. 日报存在哪里?
M9 以后日报不再生成 HTML,而是结构化写入 MySQL:
- `news_report`:日报主表(`report_type='intl'`)
- `news_event`:事件明细(`section='intl'`)
历史 HTML 日报(2026-06-16 ~ 2026-08-03)仍保留在 `data/reports/` 或已由 news 项目导入。
## 3. 如何断点续跑?
```bash
bash scripts/domestic_full.sh --resume
bash scripts/pipeline.sh --resume
```
- 步骤状态:`data/run_state/{YYYYMMDD}.state`
- 失败步骤不会标记,`--resume` 会重跑失败步骤
- 各步骤还有文件级增量,重复运行不会重复处理
## 4. 如何添加新源?
编辑 `configs/sources.yaml`,参考现有源增加一条配置,然后:
```bash
uv run en-news crawl --source <new_id>
uv run en-news extract --source <new_id>
```
详细字段见 [配置说明](configuration.md)。
## 5. 为什么 translate 或 report 报 API Key/MySQL 错误?
检查:
```bash
# 是否已加载 .env
grep -E 'DEEPSEEK|DASHSCOPE|NEWS_DB' .env | sed 's/=.*/=***/'
# DeepSeek
uv run en-news translate
# MySQL
uv run en-news report
```
CLI 入口会 `load_dotenv()`,Shell 脚本也会加载 `.env`。
## 6. Qdrant 本地文件模式和远程模式如何选择?
- 默认本地文件模式:`data/qdrant_storage`,零运维,适合 Pi 单机。
- 远程模式:`.env` 配置非 localhost 的 `QDRANT_URL` 和 `QDRANT_API_KEY`,适合多端共享。
切换后需要确保 collection 中向量由同一 embedding 模型生成。
## 7. 是否可以只跑管道不生成日报?
可以:
```bash
uv run en-news pipeline --skip-report
```
注意 `scripts/pipeline.sh` 目前固定包含日报;如果 MySQL 未配置,可用 CLI `--skip-report` 或改用 `uv run en-news translate`、`embed`、`index` 手动执行。
## 8. 为什么日志里出现“AI 大模型(场景 ...)”?
这是项目刻意保留的可观测性输出:
- M4 会打印 `AI 大模型(场景 translation)`
- 日报会打印 `AI 大模型(场景 daily_report)`
- M5 会打印 Embedding 初始化信息
方便在日志中确认实际调用的供应商/模型。
## 9. 历史 AI Agent 文档去哪了?
已清理:
- 根目录 `CLAUDE.md`
- 根目录 `continuation.md`
- 根目录 `english-news-plan.md`
- 旧版 `docs/intlnews_usage.html` / `docs/intlnews_usage.md`
所有有效内容已整理进当前 `docs/` 文档树。
## 10. 测试有失败是否影响生产?
当前已知 2 个 `test_crawler.py` 失败,属于测试路径问题,不影响生产管道。建议后续修复 `crawler/storage.py:load_index()` 对测试临时目录的适配。
-390
View File
@@ -1,390 +0,0 @@
<!DOCTYPE html>
<html lang="zh-CN">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>IntlNews 操作手册</title>
<style>
:root {
--bg: #1a1a2e;
--card: #16213e;
--accent: #0f3460;
--green: #00c853;
--yellow: #ffd600;
--red: #ff1744;
--text: #e0e0e0;
--muted: #8892b0;
}
* { margin: 0; padding: 0; box-sizing: border-box; }
body { font-family: -apple-system, 'Segoe UI', system-ui, sans-serif; background: var(--bg); color: var(--text); line-height: 1.7; padding: 20px; }
.container { max-width: 960px; margin: 0 auto; }
h1 { font-size: 2em; margin: 1em 0 0.3em; color: #fff; border-bottom: 2px solid var(--accent); padding-bottom: 0.3em; }
h2 { font-size: 1.4em; margin: 1.5em 0 0.5em; color: #64ffda; }
h3 { font-size: 1.1em; margin: 1em 0 0.3em; color: #80cbc4; }
p, li { margin: 0.4em 0; }
code { background: #0d2137; padding: 2px 6px; border-radius: 3px; font-size: 0.9em; color: #a8d8ea; }
pre { background: #0d2137; padding: 12px 16px; border-radius: 6px; overflow-x: auto; border: 1px solid #1a3a5c; margin: 0.8em 0; }
table { width: 100%; border-collapse: collapse; margin: 1em 0; }
th, td { padding: 8px 12px; text-align: left; border-bottom: 1px solid #1a3a5c; }
th { background: var(--accent); color: #fff; font-weight: 600; }
tr:hover { background: #1a2a4a; }
.tag { display: inline-block; padding: 1px 8px; border-radius: 10px; font-size: 0.8em; font-weight: 600; }
.tag-ok { background: #1b5e20; color: #a5d6a7; }
.tag-warn { background: #e65100; color: #ffcc80; }
.tag-err { background: #b71c1c; color: #ef9a9a; }
.tag-new { background: #0d47a1; color: #90caf9; }
.info-box { background: #0d2137; border-left: 4px solid #64ffda; padding: 12px 16px; border-radius: 0 6px 6px 0; margin: 1em 0; }
.warn-box { background: #1e1a0d; border-left: 4px solid #ffd600; padding: 12px 16px; border-radius: 0 6px 6px 0; margin: 1em 0; }
.cmd { display: block; padding: 4px 0; font-family: 'SF Mono', 'Fira Code', monospace; font-size: 0.9em; }
.cmd::before { content: '$ '; color: #64ffda; }
.footer { margin-top: 3em; padding-top: 1em; border-top: 1px solid #1a3a5c; color: var(--muted); font-size: 0.85em; text-align: center; }
</style>
</head>
<body>
<div class="container">
<h1>📰 IntlNews — 国际财经新闻 Deep Research</h1>
<p style="color:var(--muted);">最后更新:2026-07-23 | 运行服务器:pi@192.168.1.160</p>
<div class="info-box">
<strong>项目定位:</strong>面向国际财经新闻的私有化 Deep Research 平台。抓取英文财经新闻 → 英译中 → 事件抽取 → 向量入库 → MCP Agent 深度研究。
</div>
<!-- ==================== 一、架构 ==================== -->
<h1>一、系统架构</h1>
<img src="data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 400 250'%3E%3Crect width='400' height='250' fill='%231a1a2e'/%3E%3Ctext x='20' y='30' fill='%2364ffda' font-size='14'%3EPi 4 (8GB) — 全链路独立运行%3C/text%3E%3C/text%3E%3C/svg%3E" alt="Arch" style="display:none;">
<pre>
┌──────────────────────────────────────────────────────┐
│ Pi 4 (8GB) │
│ ┌──────────┐ ┌──────────┐ ┌──────────────────┐ │
│ │ M1 抓取 │──▶│ M2 正文 │──▶│ M3 三层去重 │ │
│ │ headful │ │ trafilat.│ │ URL/Content/SimH │ │
│ │ Playwright│ │ │ │ │ │
│ │ + HTTP 代 │ │ │ │ │ │
│ │ 理(海外) │ │ │ │ │ │
│ └──────────┘ └──────────┘ └────────┬─────────┘ │
│ ▼ │
│ ┌──────────┐ ┌──────────┐ ┌──────────────────┐ │
│ │ M6 Qdrant│◀──│ M5 向量 │◀──│ M4 翻译+事件抽取 │ │
│ │ 入库 │ │ DashScope│ │ DeepSeek v4 │ │
│ └────┬─────┘ └──────────┘ └──────────────────┘ │
│ │ │
│ ▼ │
│ ┌──────────┐ ┌──────────┐ │
│ │ M7 日报 │──▶│ doorcome │ │
│ │ 生成 │ │ 上传 │ │
│ └──────────┘ └──────────┘ │
│ │ │
│ ▼ │
│ ┌────────────────────────────────────────────────┐ │
│ │ M8 MCP 服务(Cherry Studio / Claude Code) │ │
│ └────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────┘
</pre>
<div class="info-box">
<strong>关键变更:</strong>海外服务器已于 2026-07-14 停用。全链路在 Pi 独立运行,通过 <code>Privoxy → SOCKS5 (ss-local)</code> 代理访问海外新闻网站。<br>
<strong>反爬策略:</strong>2026-07-23 更新 — 启用 Crawl4AI <code>magic</code> 全自动反检测 + Playwright <code>stealth</code> 模式。investing.com 改用 Google News RSS 代理。
</div>
<!-- ==================== 二、部署 ==================== -->
<h1>二、部署指南</h1>
<h2>2.1 环境要求</h2>
<ul>
<li>Python 3.11 + uv</li>
<li>Xvfb(headful Playwright 虚拟显示器)</li>
<li>Shadowsocks-libev + v2ray-plugin(SOCKS5 代理)</li>
<li>Privoxy(HTTP → SOCKS5 转换)</li>
</ul>
<h2>2.2 初始化安装</h2>
<pre>
# 1. 克隆仓库
git clone ssh://gitea:2222/simon/intl_news /home/pi/intlnews
cd /home/pi/intlnews
# 2. Python 环境
uv sync
# 3. 环境变量
cp .env.example .env
# 填入: DEEPSEEK_API_KEY, DASHSCOPE_API_KEY, QDRANT_URL
# 4. Xvfb 虚拟显示器
sudo apt install xvfb
Xvfb :99 -screen 0 1280x1024x24 -ac +extension RANDR &amp;
# 5. Shadowsocks + Privoxy(海外代理)
sudo apt install shadowsocks-libev privoxy
# 配置 /etc/shadowsocks-libev/config.json
# 配置 /etc/privoxy/config: forward-socks5t / 127.0.0.1:1088 .
sudo systemctl enable --now shadowsocks-libev-local
sudo systemctl enable --now privoxy
</pre>
<h2>2.3 定时调度(crontab)</h2>
<pre>
# crontab -e 添加:
0 6 * * * cd /home/pi/intlnews && bash scripts/domestic_full.sh >> logs/full.log 2>&1
0 12 * * * cd /home/pi/intlnews && bash scripts/domestic_full.sh >> logs/full.log 2>&1
0 18 * * * cd /home/pi/intlnews && bash scripts/domestic_full.sh >> logs/full.log 2>&1
0 22 * * * cd /home/pi/intlnews && bash scripts/domestic_full.sh >> logs/full.log 2>&1
@reboot Xvfb :99 -screen 0 1280x1024x24 -ac +extension RANDR &amp;
</pre>
<!-- ==================== 三、操作命令 ==================== -->
<h1>三、操作命令</h1>
<h2>3.1 全流程运行</h2>
<pre>
# 一键全流程(M1→M6→日报)
bash scripts/domestic_full.sh
</pre>
<h2>3.2 M1 抓取</h2>
<pre>
# 抓取全部 13 个源
bash scripts/domestic_crawl_8g.sh
# 单源调试
bash scripts/domestic_crawl_8g.sh investing
bash scripts/domestic_crawl_8g.sh investinglive
bash scripts/domestic_crawl_8g.sh reuters
</pre>
<h2>3.3 M2→M6 管道</h2>
<pre>
# 完整管道(需要先有 raw 数据)
bash scripts/pipeline.sh
# 或逐步执行:
PYTHONPATH=. .venv/bin/python3 -c "from scheduler.pipeline import run_pipeline; run_pipeline('$(date +%Y%m%d)')"
</pre>
<h2>3.4 日报</h2>
<pre>
# 生成当日日报
PYTHONPATH=. .venv/bin/python3 -c "from scheduler.reporter import generate_report; print(generate_report())"
# 查看日报列表
ls -lt data/reports/
</pre>
<h2>3.5 日志查看</h2>
<pre>
# 全量日志
tail -100 logs/full.log
# 筛选指定源
grep 'investing\|investinglive' logs/full.log | tail -20
</pre>
<!-- ==================== 四、新闻源状态 ==================== -->
<h1>四、新闻源状态</h1>
<table>
<tr><th>源</th><th>ID</th><th>策略</th><th>日产量</th><th>状态</th></tr>
<tr>
<td>Reuters</td><td><code>reuters</code></td><td>Google News RSS</td><td>~30 篇</td>
<td><span class="tag tag-warn">摘要</span> DataDome</td>
</tr>
<tr>
<td>CNBC</td><td><code>cnbc</code></td><td>CNBC RSS</td><td>~25 篇</td>
<td><span class="tag tag-ok">正常</span></td>
</tr>
<tr>
<td>MarketWatch</td><td><code>marketwatch</code></td><td>RSS + stealth</td><td>~25 篇</td>
<td><span class="tag tag-ok">正常</span></td>
</tr>
<tr>
<td>Financial Times</td><td><code>ft</code></td><td>FT RSS</td><td>~20 篇</td>
<td><span class="tag tag-warn">摘要</span> 付费墙</td>
</tr>
<tr>
<td>Yahoo Finance</td><td><code>yahoo_finance</code></td><td>Yahoo RSS</td><td>~30 篇</td>
<td><span class="tag tag-ok">正常</span></td>
</tr>
<tr>
<td>Investing.com</td><td><code>investing</code></td><td><strong>Google News RSS</strong> + stealth</td><td>25 篇</td>
<td><span class="tag tag-new">已修复</span> RSS 代理</td>
</tr>
<tr>
<td>Seeking Alpha</td><td><code>seekingalpha</code></td><td>RSS + stealth</td><td>~20 篇</td>
<td><span class="tag tag-ok">正常</span></td>
</tr>
<tr>
<td>Barron's</td><td><code>barrons</code></td><td>Dow Jones RSS + stealth</td><td>~20 篇</td>
<td><span class="tag tag-ok">正常</span></td>
</tr>
<tr>
<td>WSJ</td><td><code>wsj</code></td><td>Dow Jones RSS + stealth</td><td>~20 篇</td>
<td><span class="tag tag-ok">正常</span></td>
</tr>
<tr>
<td>Economist</td><td><code>economist</code></td><td>RSS + stealth</td><td>~15 篇</td>
<td><span class="tag tag-ok">正常</span></td>
</tr>
<tr>
<td style="background:#0d47a1;">InvestingLive</td><td><code>investinglive</code></td><td><strong>RSS(全文)</strong></td><td><strong>25 篇</strong></td>
<td><span class="tag tag-new">2026-07-23 新增</span></td>
</tr>
<tr>
<td>ZeroHedge</td><td><code>zerohedge</code></td><td>FeedBurner RSS</td><td>~15 篇</td>
<td><span class="tag tag-ok">正常</span></td>
</tr>
<tr style="color:var(--muted); text-decoration:line-through;">
<td>ForexLive</td><td><del><code>forexlive</code></del></td><td>—</td><td>—</td>
<td><span class="tag tag-err">已移除</span> 301→investinglive</td>
</tr>
</table>
<div class="warn-box">
<strong>⚠️ 注意:</strong>RSS 摘要源的全文依赖 headful Playwright 回退抓取,成功率取决于代理线路质量。investing.com 因 Cloudflare 反爬,全文抓取间歇性失败。
</div>
<!-- ==================== 五、数据存储 ==================== -->
<h1>五、数据存储</h1>
<pre>
/home/pi/intlnews/
├── data/
│ ├── raw/{source_id}/{yyyymmdd}/ M1 原始抓取(.html + .md + index.jsonl)
│ ├── processed/{source_id}/ M2 正文提取结果
│ ├── dedup/ M3 去重指纹库(SQLite)
│ ├── translations/{yyyymmdd}/ M4 翻译+事件抽取(JSONL)
│ ├── embeddings/{yyyymmdd}/ M5 向量文件(JSONL)
│ ├── reports/ M7 日报(HTML)
│ └── vectorstore/ M6 Qdrant 本地数据
├── configs/
│ ├── sources.yaml 新闻源定义(13 个源)
│ ├── system.yaml 系统参数
│ └── profiles/8g_headful.yaml Pi 抓取配置
├── logs/full.log 运行日志
└── prompts/ LLM Prompt 模板
</pre>
<!-- ==================== 六、MCP 服务 ==================== -->
<h1>六、MCP 服务</h1>
<p>M8 模块提供 MCP(Model Context Protocol)服务,供 Cherry Studio / Claude Code / Cursor 等客户端接入。</p>
<h2>6.1 启动</h2>
<pre>
cd /home/pi/intlnews
PYTHONPATH=. .venv/bin/python3 -m mcp_server.server
</pre>
<h2>6.2 可用工具</h2>
<table>
<tr><th>Tool</th><th>参数</th><th>功能</th></tr>
<tr><td><code>en_news_search</code></td><td>query, top_k</td><td>语义搜索知识库</td></tr>
<tr><td><code>en_news_filter</code></td><td>source_id, date, sentiment</td><td>按条件筛选事件</td></tr>
<tr><td><code>en_news_trending</code></td><td>days</td><td>获取近期热点</td></tr>
<tr><td><code>en_news_report</code></td><td>—</td><td>获取最新日报</td></tr>
<tr><td><code>en_news_daily_brief</code></td><td>—</td><td>一键生成简报</td></tr>
</table>
<!-- ==================== 七、配置参考 ==================== -->
<h1>七、配置参考</h1>
<h2>7.1 环境变量(.env)</h2>
<table>
<tr><th>变量</th><th>用途</th></tr>
<tr><td><code>DEEPSEEK_API_KEY</code></td><td>DeepSeek LLM(翻译+事件抽取)</td></tr>
<tr><td><code>DASHSCOPE_API_KEY</code></td><td>DashScope Embedding(向量生成)</td></tr>
<tr><td><code>QDRANT_URL</code></td><td>Qdrant 服务地址</td></tr>
</table>
<h2>7.2 新闻源配置(sources.yaml)</h2>
<p>每个源包含:id, name, homepage, rss_url, article_url_pattern, anti_bot_mode 等。</p>
<p>新增源时参考现有格式,启用/禁用通过 <code>enabled: true/false</code>。</p>
<h2>7.3 Profile 配置</h2>
<p>Pi 服务器使用 <code>8g_headful</code> profile(环境变量 <code>EN_NEWS_PROFILE</code>),通过 Privoxy HTTP 代理访问海外网站。</p>
<ul>
<li><code>proxy.enabled: true</code> — 启用 HTTP 代理</li>
<li><code>crawler.headful: true</code> — 有头浏览器 + Xvfb</li>
<li><code>crawler.magic: true</code> — Crawl4AI 全自动反检测</li>
<li><code>crawler.enable_stealth: true</code> — Playwright stealth</li>
</ul>
<!-- ==================== 八、排障 ==================== -->
<h1>八、常见排障</h1>
<table>
<tr><th style="width:200px;">现象</th><th>排查步骤</th></tr>
<tr>
<td>抓取 0 篇</td>
<td>
1. 检查 <code>logs/full.log</code> 看 RSS/Web 错误<br>
2. <code>bash scripts/domestic_crawl_8g.sh &lt;source_id&gt;</code> 单源测试<br>
3. 检查代理是否正常:<code>curl -I --proxy 127.0.0.1:3128 https://example.com</code>
</td>
</tr>
<tr>
<td>RSS 超时</td>
<td>
1. 确认 Shadowsocks 和 Privoxy 在运行<br>
2. <code>systemctl status shadowsocks-libev-local privoxy</code><br>
3. 环境变量 <code>HTTP_PROXY</code> 是否设置
</td>
</tr>
<tr>
<td>Cloudflare 拦截</td>
<td>
1. 已启用 <code>magic</code> + <code>stealth</code> 自动反检测<br>
2. 仍失败则需更换代理 IP(数据中心 IP 被标记)<br>
3. 改用 Google News RSS 代理(已为 investing/reuters 启用)
</td>
</tr>
<tr>
<td>日报无数据</td>
<td>
1. 检查 <code>data/translations/{date}/</code> 是否有 JSONL<br>
2. 检查 <code>data/embeddings/{date}/</code> 是否有向量
</td>
</tr>
<tr>
<td>MCP 连接失败</td>
<td>
1. 确认 MCP 服务进程在运行<br>
2. 检查端口监听:<code>ss -tlnp | grep 8080</code>
</td>
</tr>
</table>
<!-- ==================== 九、变更记录 ==================== -->
<h1>九、变更记录</h1>
<table>
<tr><th>日期</th><th>变更</th></tr>
<tr>
<td>2026-07-23</td>
<td>
<strong>反爬策略增强:</strong><br>
- crawler.py: magic/stealth/--disable-webrtc/remove_consent_popups<br>
- rss_crawler.py: httpx 代理支持<br>
- investing.com: Google News RSS 代理(绕过 Cloudflare)<br>
- <strong>forexlive → investinglive 迁移</strong>(域名已 301 重定向)
</td>
</tr>
<tr>
<td>2026-07-14</td>
<td>
- 海外服务器停用,Pi 全链路独立运行<br>
- 8g_headful profile 新增(headful + HTTP 代理)<br>
- Shadowsocks + Privoxy 部署
</td>
</tr>
</table>
<div class="footer">
IntlNews — 国际财经新闻私有化 Deep Research 平台<br>
运行于 pi@192.168.1.160 | 调度 06:00 / 12:00 / 18:00 / 22:00
</div>
</div>
</body>
</html>
-672
View File
@@ -1,672 +0,0 @@
# 国际财经 Deep Research 平台 — 使用手册
> 版本:v1.0
> 最后更新:2026-06-21
> 项目路径:国内 `/home/pi/intlnews` / 海外 `/opt/intlgrab`
---
## 目录
1. [项目概述](#1-项目概述)
2. [部署拓扑](#2-部署拓扑)
3. [环境配置](#3-环境配置)
4. [CLI 命令参考](#4-cli-命令参考)
5. [M1 — 新闻抓取](#5-m1--新闻抓取)
6. [M2 — 正文提取](#6-m2--正文提取)
7. [M3 — 三层去重](#7-m3--三层去重)
8. [M4 — 翻译 + 事件抽取](#8-m4--翻译--事件抽取)
9. [M5 — 向量生成](#9-m5--向量生成)
10. [M6 — Qdrant 入库与检索](#10-m6--qdrant-入库与检索)
11. [M7 — 全链路管道 + 日报](#11-m7--全链路管道--日报)
12. [M8 — MCP 服务](#12-m8--mcp-服务)
13. [数据目录结构](#13-数据目录结构)
14. [定时任务](#14-定时任务)
15. [常见问题](#15-常见问题)
---
## 1. 项目概述
本项目构建面向**国际英文财经新闻**的私有化 Deep Research 平台。
核心能力:
- 英文财经新闻抓取(Crawl4AI,12 个源)
- 正文提取(trafilatura)
- 全文英译中(DeepSeek LLM)
- 投资事件抽取(美股代码识别 + 情绪判断 + 重要度评分)
- 双语向量知识库(Qdrant,1024 维)
- 语义检索(中文自然语言)
- MCP 服务(Claude Code / Cherry Studio Agent 深度研究)
- 每日 AI 摘要日报(HTML)
**本项目不是**:交易系统 / 股票预测系统 / 投资顾问系统。
---
## 2. 部署拓扑
```
┌──────────────────────────────────────────────────┐
│ Overseas Server (海外) │
│ M1 Crawl4AI 抓取 → data/raw/ │
│ 每天 4 次打包 → rsync 推送 │
│ SSH: <海外服务器> │
│ 路径: /opt/intlgrab │
└────────────────────┬─────────────────────────────┘
│ rsync
▼
┌──────────────────────────────────────────────────┐
│ Domestic Server (国内) │
│ M2 正文提取 → M3 去重 → M4 翻译+事件 │
│ → M5 向量生成 → M6 Qdrant 入库 │
│ → M7 调度 + 日报 → M8 MCP 服务 │
│ SSH: <国内服务器> │
│ 路径: /home/pi/intlnews │
└──────────────────────────────────────────────────┘
```
---
## 3. 环境配置
### 3.1 依赖安装
```bash
cd /home/pi/intlnews
uv sync
```
### 3.2 配置文件
| 文件 | 用途 | 示例 |
|------|------|------|
| `.env` | API Key / URL(不入 Git) | `DEEPSEEK_API_KEY=sk-xxx` |
| `configs/system.yaml` | 功能参数(模型、阈值、超时) | `llm.provider: deepseek` |
| `configs/sources.yaml` | 新闻源定义 | 12 个英文财经源 |
### 3.3 必需环境变量(`.env`)
```bash
# DeepSeek(M4 翻译+事件抽取)
DEEPSEEK_API_KEY=sk-your-key
DEEPSEEK_BASE_URL=https://api.deepseek.com
# Qwen(备选 LLM)
QWEN_API_KEY=sk-your-key
QWEN_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
# DashScope(M5 向量生成)
DASHSCOPE_API_KEY=sk-your-key
# Qdrant(留空使用本地文件模式)
QDRANT_URL=http://localhost:6333
QDRANT_API_KEY=
```
### 3.4 关键配置项(`configs/system.yaml`)
```yaml
crawler:
max_memory_mb: 1800 # 串行抓取内存上限
dedup:
hamming_distance_threshold: 3 # SimHash 汉明距离阈值
simhash_window_days: 30 # 时间窗口
llm:
provider: "deepseek"
deepseek_model: "deepseek-v4-flash"
concurrency: 3 # LLM 并发数
embedding:
provider: "dashscope"
dashscope_model: "text-embedding-v3"
dimension: 1024
qdrant:
collection: "en_finance_news"
schedule:
day_cutoff_hour: 6 # 新闻日切分点(06:00)
```
---
## 4. CLI 命令参考
所有命令通过 `uv run en-news` 执行:
| 命令 | Milestone | 功能 |
|------|-----------|------|
| `crawl` | M1 | 抓取英文财经新闻 |
| `extract` | M2 | 正文提取 |
| `dedup` | M3 | 三层去重 |
| `translate` | M4 | 翻译 + 事件抽取 |
| `embed` | M5 | 向量生成 |
| `index` | M6 | Qdrant 入库 |
| `search <query>` | M6 | 语义检索 |
| `pipeline` | M7 | 一键全链路 M2→M6 |
| `report` | M7 | 日报生成 |
| `mcp-server` | M8 | 启动 MCP 服务 |
常用选项:
```bash
# 指定源
uv run en-news crawl --source forexlive
uv run en-news extract --source forexlive
# 指定 top_k
uv run en-news search "美联储利率决议" --top-k 5
# 重建 Qdrant collection
uv run en-news index --recreate
# 全链路跳过日报
uv run en-news pipeline --skip-report
```
---
## 5. M1 — 新闻抓取
### 5.1 手动抓取
```bash
# 抓取所有启用的源
uv run en-news crawl
# 只抓取指定源
uv run en-news crawl -s forexlive
```
### 5.2 新闻源列表
| 源 ID | 名称 | 类型 |
|-------|------|------|
| reuters | Reuters | 综合财经 |
| cnbc | CNBC | 市场新闻 |
| marketwatch | MarketWatch | 市场数据 |
| ft | Financial Times | 财经深度 |
| yahoo_finance | Yahoo Finance | 综合 |
| investing | Investing.com | 全球市场 |
| seekingalpha | Seeking Alpha | 投资分析 |
| barrons | Barrons | 市场评论 |
| wsj | WSJ | 综合财经 |
| economist | The Economist | 经济分析 |
| forexlive | ForexLive | 外汇新闻 |
| zerohedge | ZeroHedge | 另类财经 |
### 5.3 反爬策略
部分新闻源设有反爬保护(如 DataDome)。系统支持三种策略,按优先级自动选择:
| 优先级 | 策略 | 配置字段 | 说明 | 适用源 |
|--------|------|---------|------|--------|
| 1 | **RSS 抓取** | `rss_url` | 通过 RSS/Atom Feed 获取文章列表,完全绕过反爬 | MarketWatch ✅ |
| 2 | **Stealth 模式** | `anti_bot_mode: "stealth"` | 隐藏 webdriver 特征(`--disable-blink-features=AutomationControlled`) | Reuters(海外内存不足待验证) |
| 3 | **Headful 模式** | `anti_bot_mode: "headful"` | 非 headless 浏览器,最像真人 | 重度反爬源回退 |
配置示例(`configs/sources.yaml`):
```yaml
- id: "marketwatch"
rss_url: "https://feeds.marketwatch.com/marketwatch/topstories" # RSS 优先
anti_bot_mode: "headful" # RSS 失败时回退
- id: "reuters"
anti_bot_mode: "stealth" # 无 RSS,直接 stealth
```
**已验证**:
| 源 | 方式 | 结果 |
|----|------|------|
| MarketWatch | RSS | ✅ 10/10 |
| WSJ | RSS | ✅ 20/20 |
| ForexLive | 标准 headless | ✅ 17/17 |
| ZeroHedge | 标准 + URL 过滤 | ✅ 15/15 |
| Barron's | RSS (Dow Jones) | ✅ 10/10 |
| Reuters | stealth | ⚠️ DataDome |
| SeekingAlpha | stealth | ⚠️ PerimeterX |
### 5.4 海外定时抓取
```cron
# crontab(<海外服务器>)— 每天 4 次
# 时序: 抓取(60min) → 打包(5min) → 10min后国内拉取
0 6 * * * cd /opt/intlgrab && bash scripts/overseas_crawl.sh
5 7 * * * cd /opt/intlgrab && bash scripts/overseas_pack.sh
0 12 * * * cd /opt/intlgrab && bash scripts/overseas_crawl.sh
5 13 * * * cd /opt/intlgrab && bash scripts/overseas_pack.sh
0 18 * * * cd /opt/intlgrab && bash scripts/overseas_crawl.sh
5 19 * * * cd /opt/intlgrab && bash scripts/overseas_pack.sh
0 22 * * * cd /opt/intlgrab && bash scripts/overseas_crawl.sh
5 23 * * * cd /opt/intlgrab && bash scripts/overseas_pack.sh
```
### 5.4 增量抓取机制
抓取采用**两层增量**确保不重复下载和存储:
| 层级 | 位置 | 机制 |
|------|------|------|
| 抓取层 | `_extract_article_urls` | 读取当日 `index.jsonl` 中已抓取的 url_hash,跳过已存在的 URL,不重复下载 |
| 存储层 | `write_index_jsonl` | 追加写入前再次检查 url_hash,已存在则跳过 |
同一天内多次执行 `crawl`,只有新文章才会被下载和存储。
### 5.5 产物
```
data/raw/{source_id}/{YYYYMMDD}/
├── {url_hash}.html # 原始 HTML
├── {url_hash}.md # Crawl4AI Markdown
└── index.jsonl # 文章索引
```
---
## 6. M2 — 正文提取
### 6.1 执行
```bash
# 提取所有源
uv run en-news extract
# 提取指定源
uv run en-news extract -s forexlive
```
### 6.2 技术方案
- 优先使用 Crawl4AI 输出的 Markdown
- 回退 `trafilatura` 英文正文提取
- 最小正文字数阈值:50 词
### 6.3 产物
```
data/processed/{source_id}/{YYYYMMDD}/
├── {url_hash}.json # ProcessedArticle
└── index.jsonl # 处理索引
```
---
## 7. M3 — 三层去重
### 7.1 执行
```bash
uv run en-news dedup
```
### 7.2 去重逻辑
| 层级 | 方法 | 说明 |
|------|------|------|
| L1 | URL Hash | 完全相同 URL 直接命中 |
| L2 | 内容 Hash | 标准化后 SHA1[:16] 匹配(去标点/空白) |
| L3 | SimHash | 字符 3-gram,汉明距离 ≤ 3,30 天窗口 |
### 7.3 产物
```
data/deduped/{YYYYMMDD}/
├── uniques/{url_hash}.json # 唯一文章
└── index.json # 去重索引
data/dedup/fingerprints.sqlite3 # 指纹库
```
---
## 8. M4 — 翻译 + 事件抽取
### 8.1 执行
```bash
uv run en-news translate
```
### 8.2 技术方案
- **Provider**: DeepSeek v4-flash(默认)/ Qwen 备选
- **单次调用**:翻译 + 事件抽取合并,节省 token
- **并发**:3 线程(`system.yaml` → `llm.concurrency`)
- **重试**:3 次指数退避(1s → 2s → 4s)
### 8.3 输出格式
```json
{
"title": "Fed Holds Rates Steady as Markets Rally",
"title_zh": "美联储维持利率不变,市场上涨",
"content_en": "The Federal Reserve held...",
"content_zh": "美联储周三维持利率不变...",
"events": [
{
"event_type": "央行决议",
"stock_codes": [],
"sentiment": "neutral",
"importance": 5,
"summary_zh": "美联储维持利率不变,市场反弹"
}
]
}
```
### 8.4 14 种事件类型
`财报披露` `并购收购` `产品发布` `监管政策` `宏观经济`
`央行决议` `行业动态` `技术突破` `高管变动` `诉讼法律`
`市场异动` `地缘政治` `大宗商品` `外汇波动` `其他`
### 8.5 产物
```
data/events/{YYYYMMDD}/
├── {url_hash}.json # EnTranslatedArticle
└── index.json # 事件索引
```
---
## 9. M5 — 向量生成
### 9.1 执行
```bash
uv run en-news embed
```
### 9.2 技术方案
- **Provider**: DashScope `text-embedding-v3`
- **维度**: 1024
- **嵌入文本**: `标题: {title_zh}` + `事件: [{sentiment}] {event_type} 重要度{n} {summary_zh}` + `正文: {content_zh[:3000]}`
- **截断**: 4000 字符上限
### 9.3 产物
```
data/embeddings/{YYYYMMDD}/
├── {url_hash}.json # EmbeddingResult(1024 维)
└── index.json # 向量索引
```
---
## 10. M6 — Qdrant 入库与检索
### 10.1 入库
```bash
# 增量入库
uv run en-news index
# 重建 collection(清空旧数据)
uv run en-news index --recreate
```
### 10.2 语义搜索
```bash
# 基本搜索
uv run en-news search "美联储利率决议"
# 指定返回条数
uv run en-news search "伊朗霍尔木兹海峡" --top-k 5
```
### 10.3 技术方案
- **模式**: 本地文件(`data/qdrant_storage/`),无需 Docker
- **Collection**: `en_finance_news`
- **距离**: Cosine
- **Payload**: title / title_zh / url / source_id / events / content_zh_preview
### 10.4 产物
```
data/qdrant_storage/ # Qdrant 本地文件存储
```
---
## 11. M7 — 全链路管道 + 日报
### 11.1 手动全链路运行(完整流程)
当需要手工执行完整的数据处理流程时,按顺序执行以下命令:
```bash
# 步骤 1(海外): 抓取英文财经新闻
ssh <海外服务器> "cd /opt/intlgrab && uv run en-news crawl"
# 步骤 2(海外): 打包 raw 数据
ssh <海外服务器> "cd /opt/intlgrab && bash scripts/overseas_pack.sh"
# 步骤 3(国内): 拉取海外数据
cd /home/pi/intlnews && bash scripts/domestic_sync.sh
# 步骤 4: 一键全链路 M2→M6 + 日报
cd /home/pi/intlnews && uv run en-news pipeline
```
也可以分步执行(适合调试):
```bash
# 分步模式
uv run en-news extract # M2: 正文提取
uv run en-news dedup # M3: 三层去重
uv run en-news translate # M4: 翻译 + 事件抽取
uv run en-news embed # M5: 向量生成
uv run en-news index # M6: Qdrant 入库
uv run en-news report # 日报生成
```
### 11.2 一键管道(自动)
```bash
# 全链路 M2→M6 + 日报
uv run en-news pipeline
# 跳过日报生成
uv run en-news pipeline --skip-report
# 只生成日报
uv run en-news report
```
### 11.2 管道步骤
```
M2 extract → M3 dedup → M4 translate → M5 embed → M6 index → 日报
```
每步失败记录日志但不阻断后续步骤(降级继续)。
### 11.3 HTML 日报
日报包含五个板块:
1. 🤖 **AI 摘要** — LLM 根据当日 important≥4 事件生成要点总结
2. 🔥 **重要事件** — 高重要度事件表格(标题/情绪/重要度/摘要/链接)
3. 📊 **数据总览** — M1→M6 管道统计数据
4. 📈 **情绪分布** — 利好/利空/中性比例条 + 重要度分布
5. 📋 **事件类型 TOP 10**
日报输出:`data/reports/intl_news_daily_{YYYYMMDD}.html`(约 9KB),同时自动上传到
`https://echart.doorcome.cn/research/{YYYYMMDD}/intl_news_daily_{YYYYMMDD}.html`
### 11.4 国内定时调度
```cron
# crontab(<国内服务器>)— 每天 3 次
# 全流程:SSH 触发海外打包 → 下载 → 管道串行 M2→M6 → 日报
0 7 * * * cd /home/pi/intlnews && bash scripts/domestic_full.sh
0 12 * * * cd /home/pi/intlnews && bash scripts/domestic_full.sh
0 18 * * * cd /home/pi/intlnews && bash scripts/domestic_full.sh
```
`domestic_full.sh` 统一完成:远程打包 → 同步 → 管道 → 日报。
---
## 12. M8 — MCP 服务
### 12.1 启动
```bash
uv run en-news mcp-server
```
### 12.2 可用 Tool
| Tool | 功能 | 示例 |
|------|------|------|
| `search_news` | 语义检索新闻 | `search_news("美联储利率决议")` |
| `search_by_stock` | 美股代码检索 | `search_by_stock("AAPL")` |
| `search_by_sentiment` | 按情绪检索 | `search_by_sentiment("加息", sentiment="negative")` |
| `get_today_events` | 当日重要事件 | `get_today_events(importance_min=4)` |
| `get_stats` | 系统统计 | `get_stats()` |
### 12.3 Claude Code 配置
在 Claude Code 的 MCP 配置中添加:
```json
{
"mcpServers": {
"intl-news": {
"command": "uv",
"args": ["run", "en-news", "mcp-server"],
"cwd": "/home/pi/intlnews"
}
}
}
```
---
## 13. 数据目录结构
```
data/
├── raw/ # M1: 原始抓取
│ └── {source_id}/{YYYYMMDD}/
│ ├── {url_hash}.html
│ ├── {url_hash}.md
│ └── index.jsonl
├── processed/ # M2: 正文提取
│ └── {source_id}/{YYYYMMDD}/
│ └── {url_hash}.json
├── dedup/ # M3: 指纹库
│ ├── fingerprints.sqlite3
│ └── {YYYYMMDD}/
│ ├── uniques/{url_hash}.json
│ └── index.json
├── events/ # M4: 翻译+事件
│ └── {YYYYMMDD}/
│ └── {url_hash}.json
├── embeddings/ # M5: 向量
│ └── {YYYYMMDD}/
│ └── {url_hash}.json
├── qdrant_storage/ # M6: Qdrant 本地存储
└── reports/ # M7: 日报
└── intl_news_daily_{YYYYMMDD}.html
```
---
## 14. 定时任务
### 14.1 时间线
```
海外 国内
───────────────────────────── ─────────────────────────
06:00 crawl (≈60min) 07:00 全流程(打包→下载→管道→日报)
12:00 crawl (≈60min) 12:00 全流程
18:00 crawl (≈60min) 18:00 全流程
22:00 crawl (≈60min) (夜间 crawl 结果次日 07:00 处理)
```
国内 `domestic_full.sh` 流程:`SSH触发海外打包 → 下载 → M2→M3→M4→M5→M6 → 日报`(串行)。
### 14.2 新闻日定义
- 切分点:`day_cutoff_hour: 6`(凌晨 06:00)
- 当天 06:00 至次日 05:59 属于同一个新闻日
- 例如:2026-06-19 04:00 → 新闻日 "20260618"
---
## 15. 常见问题
### Q: 如何新增新闻源?
编辑 `configs/sources.yaml`,添加源配置:
```yaml
- id: "new_source"
name: "New Source Name"
enabled: true
homepage: "https://example.com/finance/"
article_url_pattern: "/news/[^/]+/"
js_render: false
max_articles_per_run: 30
```
### Q: 翻译质量不好怎么办?
1. 调整 `configs/system.yaml` 中 `llm.temperature`(降低更保守)
2. 编辑 `prompts/translation_and_extraction.md` 优化 Prompt
3. 切换 Provider:`llm.provider: "qwen"`
### Q: Qdrant 检索太慢?
- 本地文件模式已足够快(17 条 < 0.01s)
- 数据量 > 10 万条时建议切换到 Docker 模式
- 设置 `QDRANT_URL=http://your-server:6333`
### Q: 如何查看日志?
```bash
tail -f logs/sync.log # 同步日志(国内)
tail -f logs/pipeline.log # 管道日志(国内)
tail -f logs/crawl.log # 抓取日志(海外)
tail -f logs/pack.log # 打包日志(海外)
```
日志自动清理:每周日凌晨 3 点删除 14 天前的 `.log` 文件(`scripts/cleanup_logs.sh`)。
### Q: 数据如何备份?
```bash
# 备份整个 data 目录
tar czf intlnews_backup_$(date +%Y%m%d).tar.gz data/
```
---
## 附录:技术栈
| 组件 | 技术 |
|------|------|
| 语言 | Python 3.11 |
| 包管理 | uv + pyproject.toml |
| 抓取 | Crawl4AI + Playwright |
| 正文提取 | trafilatura |
| 去重 | SimHash + SQLite |
| LLM | DeepSeek v4-flash(OpenAI SDK) |
| Embedding | DashScope text-embedding-v3 |
| 向量库 | Qdrant(本地文件模式) |
| MCP | FastMCP |
| CLI | Typer |
| 配置 | YAML + .env |
| 数据模型 | Pydantic v2 |
+162
View File
@@ -0,0 +1,162 @@
# 流水线详解
本文档描述 M1→M9 每一步的输入、输出、配置与增量机制。
## M1 新闻抓取
- 模块:`crawler/`
- 输入:`configs/sources.yaml` 中的新闻源配置
- 输出:`data/raw/{source_id}/{YYYYMMDD}/`
- `index.jsonl`:本次/当日抓取索引,包含 `url_hash`、标题、URL、摘要、时间
- `{url_hash}.html` / `{url_hash}.md`:原始网页或 Markdown
- 策略:
1. 有 `rss_url` 的源先走 RSS/Atom/Google News RSS;
2. RSS 失败或返回空时,根据 `anti_bot_mode` 走 stealth/headful Web;
3. 单源串行,`index.jsonl` 按 `url_hash` 追加去重。
- 常用命令:
- `bash scripts/domestic_crawl_8g.sh [source_id]`
- `uv run en-news crawl --source cnbc`
## M2 正文提取
- 模块:`extractor/`
- 输入:`data/raw/{source_id}/{date}/index.jsonl`
- 输出:`data/processed/{source_id}/{date}/{url_hash}.json`
- `ProcessedArticle` 包含标题、正文 `content`、词数、提取器名、发布时间等。
- 提取器:`trafilatura`,失败时回退 Markdown。
- 增量:已存在同名 `{url_hash}.json` 则跳过。
```bash
uv run en-news extract
# 或只处理一个源
uv run en-news extract --source cnbc
```
## M3 三层去重
- 模块:`dedup/`
- 输入:`data/processed/**/{date}/*.json`
- 输出:
- `data/deduped/{date}/uniques/{url_hash}.json`(唯一篇)
- `data/dedup/fingerprints.sqlite3`(指纹库)
- `data/deduped/{date}/index.json`(去重统计)
- 逻辑:
1. URL 完全一致 → 重复
2. 正文 hash 一致 → 重复
3. SimHash 汉明距离 ≤ 阈值 → 模糊重复
- 跨源合并:重复篇的来源 ID 会追加到保留唯一篇的 `source_ids`,日报可展示多来源。
```bash
uv run en-news dedup
```
## M4 翻译 + 投资事件抽取
- 模块:`llm/`
- 输入:`data/deduped/{date}/uniques/*.json`
- 输出:`data/events/{date}/{url_hash}.json`
- `EnTranslatedArticle`:英文原字段 + `title_zh` / `content_zh` / `events` / `source_ids`
- 事件结构:`event_type` / `stock_codes` / `sentiment` / `importance` / `summary_zh`
- LLM:默认 DeepSeek `deepseek-v4-flash`,场景 `translation`
- 并发:默认 3 线程,来自 `system.yaml llm.concurrency`
- 增量:`events/{date}/{url_hash}.json` 存在则跳过。
```bash
uv run en-news translate
```
## M5 向量生成
- 模块:`embedding/`
- 输入:`data/events/{date}/*.json`
- 输出:`data/embeddings/{date}/{url_hash}.json`
- 包含 `url_hash`、`source_id`、`vector` 等
- `data/embeddings/{date}/index.json` 记录该批统计
- 向量:DashScope `text-embedding-v3`,默认 1024 维、batch_size=10
- 增量:已存在向量文件则跳过。
```bash
uv run en-news embed
```
## M6 Qdrant 入库与检索
- 模块:`vectorstore/`
- 输入:`data/embeddings/{date}/` 与 `data/events/{date}/`
- 输出:Qdrant collection `en_finance_news`
- collection 大小:1024 维,余弦距离
- point ID:`url_hash` 通过 UUIDv5 转换,确定性幂等
- payload:标题、双语标题、事件、来源、时间、正文预览等
- 支持模式:
- 本地文件模式(默认):`data/qdrant_storage`
- 远程 HTTP 模式:`.env` 中配置非本地的 `QDRANT_URL`
- 命令:
```bash
# 入库
uv run en-news index
# 重建 collection(会清空)
uv run en-news index --recreate
# 重建并全量回灌所有历史向量
uv run en-news index --recreate --all
# 指定日期入库
uv run en-news index --date 20260801
# 检索
uv run en-news search "苹果 财报"
```
## M7 全链路管道编排
- 模块:`scheduler/pipeline.py`
- 串行执行:`extract → dedup → translate → embed → index → report`
- 每步失败不阻断后续,最终返回各步骤成功/失败统计。
- CLI:`uv run en-news pipeline` / `--skip-report`
- Shell:`bash scripts/pipeline.sh [--resume]`
## M8 MCP 服务
- 模块:`mcp_server/server.py`
- 启动:`uv run en-news mcp-server`
- 工具:
- `search_news`:语义搜索
- `search_by_stock`:按美股代码搜索
- `search_by_sentiment`:按情绪过滤搜索
- `get_today_events`:当日重要事件
- `get_stats`:系统统计
- 用途:供 Claude Code / Cherry Studio 等 MCP 客户端做深度研究。
## M9 日报结构化入库
- 模块:`scheduler/reporter.py` + `report_db/`
- 时间窗口:最近 25 小时
- 输出:MySQL `myquant` 库
- `news_report`:`report_type='intl'`、`report_date`、`generated_at`、`ai_summary`、`stats`
- `news_event`:`section='intl'`、Top 事件明细
- 幂等:`(report_date, report_type='intl', file_name='')` 唯一键覆盖
- 命令:`uv run en-news report`
## 数据目录速查
```text
data/
├── raw/{source_id}/{YYYYMMDD}/ M1 原始
├── processed/{source_id}/{YYYYMMDD}/ M2 正文
├── dedup/fingerprints.sqlite3 M3 指纹库
├── deduped/{YYYYMMDD}/uniques/ M3 去重后唯一篇
├── events/{YYYYMMDD}/ M4 翻译+事件
├── embeddings/{YYYYMMDD}/ M5 向量
├── qdrant_storage/ M6 本地 Qdrant
├── run_state/ 步骤级 --resume 状态
└── reports/ 历史 HTML 日报(M9 已不再产出)
```
## 失败恢复建议
1. 如果某个 Python 步骤失败,直接重跑同一命令即可,已完成的文件会自动跳过。
2. 如果 Shell 脚本中断,使用 `bash scripts/domestic_full.sh --resume` 或 `bash scripts/pipeline.sh --resume`。
3. 日报失败通常与 MySQL 连接/凭据有关,可先单独执行 `uv run en-news report` 查看日志。
4. 若 Qdrant 本地文件损坏,可考虑备份后删除 `data/qdrant_storage` 并重新 `uv run en-news index`(需重新入库全部向量)。
+93
View File
@@ -0,0 +1,93 @@
# 快速开始
## 1. 环境要求
- Python 3.11(项目支持 `>=3.11,<3.13`)
- [uv](https://docs.astral.sh/uv/) 包管理器
- Linux(推荐 Pi 4 / 8GB 内存或更高)
- 可选但推荐:Playwright 浏览器依赖、Xvfb、Privoxy/Shadowsocks(用于海外访问)
## 2. 安装依赖
```bash
# 在项目根目录执行
uv sync
# 如需开发依赖(pytest、ruff)
uv sync --extra dev
```
## 3. 配置环境变量
```bash
cp .env.example .env
```
然后编辑 `.env`,至少配置以下密钥:
| 变量 | 用途 |
|------|------|
| `DEEPSEEK_API_KEY` | M4 翻译+事件抽取、M7 日报摘要(DeepSeek) |
| `QWEN_API_KEY` 或 `DASHSCOPE_API_KEY` | Qwen LLM、M5 向量化(Embedding 复用兼容端点) |
| `DASHSCOPE_API_KEY` | M5 `text-embedding-v3` 向量化 |
| `QDRANT_URL` / `QDRANT_API_KEY` | 远程 Qdrant(可选;默认本地文件模式可留空) |
| `NEWS_DB_*` | M9 日报入库 MySQL(`NEWS_DB_PASSWORD` 必填) |
完整变量说明见 [配置说明](configuration.md)。
## 4. 验证安装
```bash
uv run en-news --help
```
能看到 `crawl / extract / dedup / translate / embed / index / search / report / pipeline / mcp-server` 即安装成功。
## 5. 首次运行
### 5.1 抓取(M1)
```bash
# 抓取全部源(8G headful + HTTP 代理,适合国内 Pi)
bash scripts/domestic_crawl_8g.sh
# 只抓取某个源
bash scripts/domestic_crawl_8g.sh reuters
```
> 如果不需要真实抓取,可先用测试数据或已有 `data/raw/` 数据跳过此步。
### 5.2 全链路管道(M2→M6 + 日报)
```bash
bash scripts/pipeline.sh
```
或使用命令行入口:
```bash
uv run en-news pipeline --skip-report # 只跑 M2→M6,不生成日报
uv run en-news pipeline # M2→M6 + 日报
```
### 5.3 语义搜索
```bash
uv run en-news search "美联储 利率 决议"
```
## 6. 日常全流程
```bash
# 全流程:M1 抓取 + 管道 + 日报
bash scripts/domestic_full.sh
# 中断后继续(跳过当天已完成步骤)
bash scripts/domestic_full.sh --resume
```
## 7. 常见首个坑
- 如果没有配置 `NEWS_DB_PASSWORD`,日报(`report`)步骤会失败。`scripts/pipeline.sh` 会包含日报;若暂时不配 MySQL,请使用 `uv run en-news pipeline --skip-report` 跳过日报。
- 抓取海外站点需要代理和 Xvfb,详见 [部署与运维](deployment.md)。
- 运行前请确保已在项目根目录(所有路径都相对项目根目录)。
+141
View File
@@ -0,0 +1,141 @@
# 使用手册
## 1. CLI 命令(`uv run en-news`)
| 命令 | 功能 | 常用参数 |
|------|------|----------|
| `crawl` | M1 抓取新闻 | `--source/-s <id>` 只抓单源;`--profile/-p` 指定 Profile |
| `extract` | M2 正文提取 | `--source/-s <id>` 只处理单源;`--date/-d` 指定日期 |
| `dedup` | M3 三层去重 | `--date/-d` 指定日期 |
| `translate` | M4 翻译+事件抽取 | `--date/-d` 指定日期 |
| `embed` | M5 向量生成 | `--date/-d` 指定日期;`--all` 处理全部日期 |
| `index` | M6 写入 Qdrant | `--recreate` 重建;`--date/-d` 指定日期;`--all` 全量回灌 |
| `search` | M6 语义检索 | 必需 query;`--top-k` 返回条数 |
| `report` | M9 生成日报并入库 | 无 |
| `pipeline` | M2→M6(+日报) | `--skip-report` 跳过日报;`--date/-d` 指定日期 |
| `mcp-server` | M8 启动 MCP 服务 | 无 |
示例:
```bash
# 只抓取 CNBC
uv run en-news crawl --source cnbc
# 单步正文提取
uv run en-news extract
# 搜索
uv run en-news search "美联储" --top-k 5
# 重建 Qdrant collection(慎用!会清空现有向量)
uv run en-news index --recreate
# 指定历史日期处理
uv run en-news translate --date 20260801
uv run en-news embed --date 20260801
uv run en-news index --date 20260801
# 全量回灌历史向量(--recreate 搭配 --all 时只会在首个日期重建 collection)
uv run en-news index --all
uv run en-news embed --all
```
## 2. Shell 脚本
| 脚本 | 用途 |
|------|------|
| `scripts/domestic_full.sh` | 国内服务器全流程:M1 抓取 + M2→M6 管道 + 日报 |
| `scripts/domestic_full.sh --resume` | 同上,但跳过当天已完成步骤 |
| `scripts/pipeline.sh` | 仅 M2→M6 管道 + 日报 |
| `scripts/pipeline.sh --resume` | 同上,支持步骤级断点续跑 |
| `scripts/domestic_crawl_8g.sh [source]` | 8G headful + HTTP 代理抓取 M1 |
| `scripts/domestic_crawl_2g.sh [source]` | 2G headless stealth 轻量抓取 M1 |
| `scripts/cleanup_logs.sh` | 清理 14 天前日志 |
| `scripts/domestic_sync.sh [date]` | 从海外服务器同步 raw 数据(海外已停用,保留兼容) |
### 2.1 全流程示例
```bash
cd /home/pi/intlnews
# 每天定时任务可执行
bash scripts/domestic_full.sh
# 手动中断后继续
bash scripts/domestic_full.sh --resume
```
### 2.2 步骤级断点说明
步骤状态记录在 `data/run_state/{YYYYMMDD}.state`:
- 每成功一个步骤追加一行,如 `M1_crawl`、`M2_extract`、`M3_dedup`、`M4_translate`、`M5_embed`、`M6_index`、`report`。
- `--resume` 只跳过已标记步骤;失败步骤不会标记,所以会从失败处重试。
- 跨天自动失效,状态文件按天隔离。
## 3. MCP 服务(M8)
MCP 是给 Cherry Studio / Claude Code 等 Agent 客户端接入的研究工具层。
启动:
```bash
uv run en-news mcp-server
```
提供工具:
| 工具 | 功能 |
|------|------|
| `search_news` | 语义检索新闻 |
| `search_by_stock` | 按美股代码检索相关事件 |
| `search_by_sentiment` | 按情绪倾向(利好/利空/中性)过滤检索 |
| `get_today_events` | 获取当日重要投资事件 |
| `get_stats` | 获取系统统计概览 |
具体接入方式由客户端决定,通常是在 MCP 配置中指向 `uv run --directory /path/to/intlnews en-news mcp-server`。
## 4. 日报查看
日报已结构化写入 MySQL:
- 主表:`news_report`
- `report_type = 'intl'`
- `ai_summary` 为 AI 日报摘要
- `stats` 为 JSON 统计快照
- 明细表:`news_event`
- `section = 'intl'`
- 包含标题、摘要、来源、情绪、重要度、URL
可通过 SQL 查看最新日报:
```sql
SELECT id, report_date, generated_at, ai_summary
FROM news_report
WHERE report_type = 'intl'
ORDER BY report_date DESC, generated_at DESC
LIMIT 5;
```
## 5. 日志
- 全流程日志:`logs/domestic_full_{YYYYMMDD_HHMMSS}.log`
- 管道日志:`logs/pipeline_{YYYYMMDD_HHMMSS}.log`
- 运行时也实时输出到终端。
快速查看最近日志:
```bash
tail -100 logs/pipeline_$(ls -t logs/pipeline_*.log | head -1 | sed 's#logs/##')
```
## 6. 常见操作组合
| 目标 | 命令 |
|------|------|
| 只抓取新闻 | `bash scripts/domestic_crawl_8g.sh` |
| 从已有 raw 跑全管道 | `uv run en-news pipeline --skip-report` |
| 增量补齐翻译 | `uv run en-news translate`(自动跳过已有) |
| 只生成日报 | `uv run en-news report` |
| 搜索知识库 | `uv run en-news search "关键词"` |
| 启动 MCP | `uv run en-news mcp-server` |