feat: AI 模型按场景独立配置(llm_scenes)+ 去重多来源合并

- system.yaml 新增 llm_scenes(translation / daily_report,含用途/方法/模型要求说明)
- load_llm_config(scene=...) 场景覆盖;日报摘要 temperature 0.3 硬编码 → 配置
- M3 去重:唯一篇记录 source_ids(跨源重复合并,首个来源为 source_id)
- source_ids 经翻译透传至 events,日报事件 source 多来源拼接展示(≤3 个)
- 已部署 pi5:merge 实证 10 源合并;日报 report_id=224 正常入库
This commit is contained in:
2026-08-12 09:56:55 +08:00
parent 9aa44e610d
commit c72a5ed13a
13 changed files with 331 additions and 20 deletions
+17 -1
View File
@@ -140,11 +140,27 @@ ls -lt data/reports/
| 文件 | 用途 |
|------|------|
| `configs/sources.yaml` | 英文财经新闻源定义(13 个源) |
| `configs/system.yaml` | 系统级业务参数(超时/并发/LLM 模型等) |
| `configs/system.yaml` | 模型按场景(`llm_scenes`)/重试/阈值等系统配置 |
| `configs/profiles/8g_headful.yaml` | Pi 服务器 headful 抓取配置(代理/超时) |
| `configs/profiles/8g_headful.yaml` | Pi 服务器 headful 抓取配置(代理/超时) |
| `.env` | 密钥 / 服务地址(不入 Git) |
| `prompts/` | LLM Prompt 模板(翻译/日报/搜索 Agent) |
### AI 模型按场景配置(llm_scenes)
大模型按场景独立配置,见 `configs/system.yaml` 的 `llm_scenes` 段:
| 场景 | 用途 | 模型(当前) | 参数 |
|------|------|-------------|------|
| `translation` | M4 全文英译中 + 投资事件抽取 | deepseek-v4-flash | temperature=0.1, max_tokens=8192 |
| `daily_report` | M7 日报 AI 摘要(分批生成) | deepseek-v4-flash | temperature=0.3, max_tokens=1500 |
场景未声明的字段回退 `llm` 默认段;Embedding 为单一场景(`en_finance_news` 库入库/检索向量必须同模型,不支持拆分)。
### 去重多来源(M3)
去重时跨源重复的新闻,会把所有来源记录到保留的唯一篇 `source_ids` 字段(首个来源为 `source_id`),经翻译透传后在日报事件 `source` 展示(如 "Barron's, CNBC, Reuters",最多 3 个)。
### 日报入库(M9)
日报内容结构化写入与 [news 项目](https://github.com/) 共用的 MySQL `myquant` 库(表 `news_report` / `news_event`,`report_type="intl"`,同一天重复生成幂等覆盖)。表结构与数据契约见 news 项目 `docs/db_schema.md`。