Files
intl_news/continuation.md
T
simon 012cb615d3 feat: 脚本执行显性输出当前阶段与 AI 模型供应商/名称
- pipeline.sh 每阶段 banner(━━━ M4 翻译+事件抽取(AI 大模型: deepseek / deepseek-v4-flash)━━━)
- ai_model_info() 从 system.yaml 读取场景模型(translation/daily_report/embedding)
- Python 层:M4/M5/日报 显性打印 provider/model(场景标注)
- 已同步 pi5 实测:客户端初始化日志含供应商/模型
2026-08-12 11:20:51 +08:00

226 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# continuation.md — English Financial News 项目状态
> 最后更新:2026-08-12
---
## 2026-08-12 会话成果(四):脚本阶段/模型显性输出
**背景:** 执行全流程时终端未显性标注"当前阶段",AI 调用的供应商/模型只在客户端初始化时输出一次。
**改动:**
| 文件 | 改动内容 |
|------|---------|
| `scripts/pipeline.sh` | 每阶段 banner(`━━━ M4 翻译+事件抽取(AI 大模型: deepseek / deepseek-v4-flash)━━━`);`ai_model_info()` 从 system.yaml 读取场景模型(translation/daily_report/embedding) |
| `scripts/domestic_full.sh` | M1 / 管道阶段标题统一 banner 风格 |
| `llm/pipeline.py` | 创建客户端后显性输出 `AI 大模型(场景 translation): provider/model` |
| `embedding/pipeline.py` | 向量化日志补充 provider |
| `scheduler/reporter.py` | daily_report 场景客户端初始化时输出 provider/model |
**测试:** 本地模拟 banner 全部正确(M4/M5/日报 3 处 AI 标注);pi5 实测客户端初始化日志含 provider/model(deepseek/deepseek-v4-flash、dashscope/text-embedding-v3);全量 184 passed。
---
## 2026-08-12 会话成果(三):全流程终端进度输出
**背景:** `domestic_full.sh` / `pipeline.sh` 原用 `tail -5/-10` 截断输出,终端看不到中间进度。
**改动:**
| 文件 | 改动内容 |
|------|---------|
| `scripts/_step_state.sh` | `step_run` 增加耗时统计与进度日志(`▶ 开始` / `✔ 完成(耗时 Ns)` / `✗ 失败(exit=N,耗时)`);支持 `STEP_LOG` 环境变量——步骤全部输出实时显示终端并同时写入日志(用 `PIPESTATUS[0]` 取原命令退出码,规避 pipefail/tee 吞退出码) |
| `scripts/pipeline.sh` | 步骤输出透传终端(去掉 `tail -10`);日志写 `logs/pipeline_{ts}.log` |
| `scripts/domestic_full.sh` | M1 与管道输出透传终端(去掉 `tail -5/-10`);日志写 `logs/domestic_full_{ts}.log` |
**测试:** 无 pipefail / pipefail+条件 / pipefail+非条件 三环境下退出码捕获均正确(失败不 mark);pi5 实机 `--resume` 跑通(M3 31s、M6 20s 均显示耗时,日志文件生成)。
**注意:** 所有含 `${var}` 后跟全角标点的输出已用花括号包裹(bash 会把全角字符字节并入变量名,导致 unbound variable)。
---
## 2026-08-12 会话成果(二):全流程 --resume 断点续跑
**背景:** `scripts/domestic_full.sh`(M1→M6→日报)中断(ssh 断线/断电)后只能从头重跑。
**改动:**
| 文件 | 改动内容 |
|------|---------|
| `scripts/_step_state.sh`(新增) | 步骤状态库:`step_run` / `step_should_run` / `step_mark` / `step_done`;状态文件 `data/run_state/{YYYYMMDD}.state`(按天换新);`step_run` 显式检查退出码(规避 bash set -e 条件上下文陷阱) |
| `scripts/pipeline.sh` | 支持 `--resume`;M2→M6→report 每步 `step_run` 包裹 |
| `scripts/domestic_full.sh` | 支持 `--resume`;M1 用 `step_should_run`/`step_mark`(部分源失败仍标记,原语义);透传 `--resume` 给 pipeline.sh |
| `.gitignore` | 忽略 `data/run_state/` |
**文件级增量确认(各步骤已内置,无需新写):** M2 processed JSON 存在跳过、M3 指纹库判重、M4 events JSON 存在跳过、M5 embedding index、M6 Qdrant upsert 幂等(id=url_hash)、日报 MySQL 唯一键覆盖。
**测试:** 本地模拟 3 场景全过(失败步骤不 mark / set -e 退出不 mark / resume 跳过已完成);`--badarg` exit=1;pi5 同步后验证 step_run 机制与参数校验正常。
**注意:** crontab 07/12/18 跑的是不带 --resume 的全量(步骤级不跳过,靠文件级增量),--resume 供手动中断恢复。
---
## 2026-08-12 会话成果
### M9.2:AI 模型按场景配置 + 去重多来源
**背景:** ① 本项目所有 AI 大模型调用点(M4 翻译+事件抽取、M7 日报 AI 摘要、M5 向量化)此前共用同一套 `llm`/`embedding` 配置;② M3 去重跨源重复时丢弃重复篇来源信息。
**改动:**
| 文件 | 改动内容 |
|------|---------|
| `configs/system.yaml` | 新增 `llm_scenes` 段(translation / daily_report,含用途/使用方法/模型要求说明,字段回退 `llm` 默认段);`embedding` 段注释说明单一场景原因 |
| `llm/client.py` | `_load_system_config(scene)` 场景合并;`load_llm_config(..., scene)` 支持按场景覆盖 provider/model/参数;场景统一用 `model` 键 |
| `llm/pipeline.py` | `load_llm_config(scene="translation")` |
| `scheduler/reporter.py` | `_call_llm_simple` 用 `scene="daily_report"`;`temperature` 硬编码 0.3 → `config.temperature`(技术债清除);新增 `_article_source_label()` 多来源拼接展示 |
| `extractor/models.py` / `llm/models.py` | `ProcessedArticle` / `EnTranslatedArticle` 新增 `source_ids` 字段 |
| `dedup/pipeline.py` | unique 初始化 `source_ids=[source_id]`;dup 时 `_merge_duplicate_source()` 跨日期目录合并来源进唯一篇 |
| `llm/extractor.py` | 同步/异步构造 `EnTranslatedArticle` 透传 `source_ids` |
| `tests/` | `test_llm.py` 场景配置 4 用例;`test_dedup.py::TestMergeSources` 3 用例;`test_report_db.py` 多来源拼接 2 用例 |
**测试结果:** 本地全量 184 passed / 2 failed(原有 test_crawler 路径问题);ruff 无新增。
**部署验证(pi5 实盘):**
- scene 配置生效:translation=(v4-flash, 0.1)、daily_report=(v4-flash, 0.3, 1500)
- 去重合并实证:历史唯一篇 `049da7e0c87160dc`(barrons 主源)被 9 个跨源重复篇合并 → `source_ids` 10 个来源
- 日报仍正常入库:report_id=224(2026-08-12, 15 事件);当天 DeepSeek 摘要 3 次返回空 → 规则兜底,日报未中断(失败处理按设计工作)
- 多来源展示:`_article_source_label` 实测 "Barron's, CNBC, Reuters";单源回退正常
**已知说明:**
- 历史 3368 个 deduped 旧文件无 `source_ids` 字段(不回填,向前生效);events 文件由下次 crontab 07:00 用新代码自然带出透传
- 179 篇全重复(增量正常);138 个跨源重复的 merge 目标多为历史日期目录文件(指纹库 ±30 天窗口所致),非当天 uniques
---
## 2026-08-04 会话成果
### M9:日报结构化入库(HTML → MySQL)
**背景:** 与 news 项目对齐,日报内容不再产出 HTML / scp 上传 doorcome,改为写入共用 MySQL `myquant` 库(表 `news_report` / `news_event`,`report_type="intl"`)。历史日报(2026-06-16 ~ 2026-08-03,intl 128 份)已由 news 项目解析入库,本次仅做新日报按相同契约入库。
**改动:**
| 文件 | 改动内容 |
|------|---------|
| `report_db/`(新增) | 复用 news 项目实现:`models.py`(EventRow/ReportData)、`schema.py`(news_report/news_event DDL)、`db.py`(load_db_config/connect/init_schema/save_report/fetch_report/exists_report,loguru 换 logging) |
| `scheduler/reporter.py` | 新增 `_build_report_data()`(intl 结构:事件全部 `section="intl"`,`file_name=""` 幂等覆盖,stats 含 pipeline/sentiment/importance/event_types/source_dist);`generate_report()` 返回 `int | None`,收集逻辑不变,末尾 `connect()+save_report()`;HTML 渲染/上传函数保留标 deprecated |
| `scheduler/pipeline.py` | `run_step_report` 适配 `report_id` 返回值 |
| `app/cli.py` | `report` 命令打印 `report_id` |
| `pyproject.toml` | `uv add pymysql`;`report_db/schema.py` 加 E501 per-file-ignore |
| `.env.example` / `README.md` | 新增 `NEWS_DB_*` 配置说明(pi5 经内网直连 `192.168.1.10:13306`,password 与 news 项目相同) |
| `tests/test_report_db.py`(新增) | 模型默认值 + `_build_report_data` 组装(板块/rank、字段映射、标题截断、stats 快照、`"?"` 归一化) |
**测试结果:** `uv run pytest` → 174 passed / 2 failed(2 个失败为原有 `test_crawler.py` index.jsonl 路径问题,与本次无关);ruff 改动文件干净(仅剩原有 `_SENTENCE_END` N806)。
**部署:** 已 rsync 至 pi5(`/home/pi/intlnews`),`.env` 写入 NEWS_DB_*,`uv sync` 安装 pymysql,真实生成日报并回读 DB 验证(详见下方)。
---
## 2026-07-23 会话成果
### investing.com 反爬策略调整
**问题:** RSS 返回 403(Cloudflare),Web headful 间歇性失败(数据中心 IP 被标记)
**改动:**
| 文件 | 改动内容 |
|------|---------|
| `crawler/crawler.py` | `magic=True` + `enable_stealth=True` Crawl4AI/Playwright 全自动反检测;`--disable-webrtc` 防止 IP 泄露;`delay_before_return_html=3s` 等待 Cloudflare 挑战;`remove_consent_popups=True` 自动关弹窗;`wait_until="domcontentloaded"` 避免被 Cloudflare 阻塞;增加列表页 URL 过滤(`most-popular-news`/`trending` 等) |
| `crawler/rss_crawler.py` | 新增 `_make_rss_client()` 读取 system.yaml 代理配置创建 httpx 客户端;Google News RSS 跳过 `article_url_pattern` 域名过滤 |
| `configs/sources.yaml` | investing.com:`rss_url` 切换为 Google News RSS 代理(绕过 Cloudflare),增加 `anti_bot_mode: stealth`;forexlive → investinglive 迁移 |
| `scripts/domestic_crawl_8g.sh` | `HTTP_PROXY`/`HTTPS_PROXY` 移至 `.env` 加载之前,确保 httpx 继承代理 |
**测试结果(investing.com):**
| 策略 | 结果 |
|------|------|
| Google News RSS(代理) | ✅ 25/25 篇,423 words/篇平均 |
| Web headful 首页 | ✅ 可加载,提取 ~25 个文章链接(15s) |
| Web headful 文章页 | ❌ 所有文章页仍被 Cloudflare 拦截(403) |
### forexlive → investinglive 迁移
**问题:** `forexlive.com` 域名 301 重定向到 `investinglive.com`,`_is_valid_article_url` 域名过滤拒绝所有链接
**改动:** `sources.yaml` — 移除 `forexlive` 源,新增 `investinglive` 源
| 字段 | 旧值 | 新值 |
|------|------|------|
| id | `forexlive` | `investinglive` |
| name | ForexLive | InvestingLive |
| homepage | `forexlive.com/` | `investinglive.com/` |
| rss_url | `forexlive.com/feed` | `investinglive.com/feed` |
| article_url_pattern | `/(news\|technical-analysis\|Education)/` | `/(news\|central-banks\|commodities\|stocks\|technical-analysis)/` |
| max_articles | 20 | 25 |
**测试结果:** ✅ `found=25 success=25`,100%,全文 3-9KB/篇(RSS 内嵌完整正文)
---
## 当前进度
```
M0 ✅ 项目骨架
M1 ✅ 新闻抓取(13/13 源,Pi headful Playwright + HTTP 代理)
M2 ✅ 正文提取(trafilatura + MD 回退,增量跳过已处理)
M3 ✅ 三层去重(L1 URL / L2 Content / L3 SimHash,SQLite 指纹库)
M4 ✅ 翻译+事件抽取(DeepSeek v4-flash,事件去重 Prompt 约束)
M5 ✅ 向量生成(DashScope text-embedding-v3,1024 维)
M6 ✅ Qdrant 入库(本地文件模式,语义搜索验证)
M7 ✅ 全链路管道 + 日报(事件级去重 + AI 摘要分批合并 + 防截断)
M8 ✅ MCP 服务(FastMCP 5 个 Tool)
```
---
## 各源抓取状态
| 源 | 策略 | 状态 | 说明 |
|----|------|------|------|
| Reuters | Google News RSS | ⚠️ 摘要,RSS 稳定 | DataDome 反爬无法突破 |
| CNBC | CNBC RSS | ✅ 稳定 | 原生 RSS 正常 |
| MarketWatch | RSS + stealth | ✅ 稳定 | RSS 优先,stealth 回退 |
| FT | FT RSS | ⚠️ 摘要,RSS 稳定 | 付费墙,全文不可用 |
| Yahoo Finance | Yahoo RSS | ✅ 稳定 | RSS 正常 |
| **Investing.com** | **Google News RSS** + stealth | **✅ RSS 25/篇,Web ❌** | **2026-07-23 修复:RSS 改用 Google News 代理** |
| Seeking Alpha | RSS + stealth | ✅ 稳定 | RSS 稳定 |
| Barron's | Dow Jones RSS + stealth | ✅ 稳定 | 无原生 RSS,Dow Jones 兜底 |
| WSJ | Dow Jones RSS + stealth | ✅ 稳定 | RSS 正常 |
| Economist | RSS + stealth | ✅ 稳定 | RSS 正常 |
| **InvestingLive** | **RSS** | **✅ 25/篇 100%** | **2026-07-23 新增(forexlive 迁移)** |
| ZeroHedge | FeedBurner RSS | ✅ 稳定 | RSS 含全文 |
| ~~Forexlive~~ | ~~RSS 已失效~~ | ~~❌ 301→investinglive~~ | **已替换为 InvestingLive** |
---
## 待办事项
1. **Investing.com 全文抓取** — Google News RSS 仅摘要,Web headful 仍被 Cloudflare 拦截(需要更好代理线路或住宅 IP)
2. **Reuters 全文** — 同上,DataDome 需要更先进的代理
3. **FT 付费墙** — 需要付费订阅或放弃全文
4. **海外服务器 crontab 停用** — 确认不再参与流水线后清理
5. **Pi crontab 确认** — 当前 `logs/full.log` 显示调度正常运行(06/12/18/22 四个时间点)
---
## 服务器状态
| 角色 | 地址 | 路径 | 状态 |
|------|------|------|------|
| 海外 | `ecs-user@8.217.19.253` | `/opt/intlgrab` | ⏸️ 不再参与流水线 |
| 国内 | `pi@192.168.1.160` | `/home/pi/intlnews` | ✅ 全链路独立运行 |
---
## CLI 命令速查
```
bash scripts/domestic_full.sh 全流程(M1→M6→日报)
bash scripts/domestic_crawl_8g.sh 仅 M1 headful(可跟源 ID)
bash scripts/domestic_crawl_8g.sh investing 单源测试
PYTHONPATH=. .venv/bin/python3 -c "..." 直接调用 orchestrator
```