Files
news/docs/architecture.md
T
simon ff911cf6f7 feat: Token Plan 迁移与 .env 热加载,并修复日报 AI 摘要为空
Token Plan 迁移 / 配置热加载:
- configs/llm_models.yaml: 各场景切到 Token Plan(deepseek-v4.1-flash / qwen3.6-flash)
- 新增 configs/runtime_env.py: .env 按 (mtime_ns, size) 热加载并同步 os.environ,
  统一 env_get 取值;llm / embedding / vectorstore / mcp / pipeline 改用 env_get
- configs/loader.py / scripts/run_scheduler.py 等配套调整
- 新增 tests/test_hot_reload.py

日报 AI 摘要为空修复(2026-09-25):
- 根因: 推理模型的 reasoning token 与正文共用 max_tokens, 预算 1500 被"思考"
  占满 -> text_tokens=0 / finish_reason=length, 摘要静默为空且不重试
- daily_report 场景新增 max_tokens(默认 4000, YAML 保存即热生效);
  LLMConfig 支持可选 max_tokens; 分块预算 800 -> 2000
- _llm_call 拆出 _call_once, 正文为空时自动加倍预算重试(上限 16000),
  用尽才降级返回空串; 网络异常重试语义不变
- docs/user-guide.md 新增 FAQ; continuation.md 记录本次排查
- 已重跑 2026-09-25 日报(report_id=357)补回 466 字摘要

测试: 相关用例 56 passed(test_hot_reload 12 passed);
      ruff 无新增问题; 3 个 crawler 既有失败与本改动无关
2026-09-25 11:13:37 +08:00

25 KiB
Raw Blame History

项目架构文档

版本:v2.0 | 最后更新:2026-08-22


目录

  1. 整体架构
  2. 数据流全景
  3. 包职责与导出 API
  4. 核心数据模型
  5. 配置体系
  6. 产物目录结构
  7. 运行环境

1. 整体架构

┌─────────────────────────────────────────────────────────────────┐
│                      统一 CLI (a_share_cli)                       │
│  crawl | extract | dedup | events | embed | ingest | pipeline    │
│  search | status | report | cninfo | watchlist | discover        │
└──────────────────────────┬──────────────────────────────────────┘
                           │
    ┌──────────────────────┼──────────────────────────────────┐
    │                      │                                  │
    ▼                      ▼                                  ▼
┌─────────┐   ┌──────────┐   ┌────────┐   ┌──────────┐   ┌──────────┐
│ crawler │──▶│extractor │──▶│ dedup  │──▶│   llm    │──▶│embedding │
│  (M1)   │   │  (M2)    │   │  (M3)  │   │  (M4)    │   │  (M5)    │
└─────────┘   └──────────┘   └────────┘   └──────────┘   └─────┬────┘
                                                                │
                                                                ▼
                                                          ┌──────────┐
                                                          │vectorstore│
                                                          │  (M6)     │
                                                          └─────┬────┘
                                                                │
                         ┌──────────────────────────────────────┤
                         │                                      │
                         ▼                                      ▼
                   ┌──────────┐                          ┌──────────┐
                   │  mcp_server│                         │scheduler │
                   │   (M8)    │                         │  (M7)    │
                   └──────────┘                          └─────┬────┘
                                                               │
                                             ┌─────────────────┤
                                             ▼                 ▼
                                       ┌──────────┐    ┌──────────────┐
                                       │ reporter │    │stock_reporter │
                                       │ (日报)    │    │  (个股日报)    │
                                       └─────┬────┘    └──────────────┘
                                             │
                              ┌──────────────┼──────────────┐
                              ▼              ▼              ▼
                        ┌──────────┐  ┌──────────┐  ┌──────────┐
                        │report_db │  │report_   │  │  MySQL   │
                        │  (M10)   │  │ import   │  │(myquant) │
                        └──────────┘  └──────────┘  └──────────┘

11 个包(按管道顺序):

序号 包 模块 概述
1 crawler/ M1 新闻抓取:httpx 静态直连 + Playwright JS 渲染 + cninfo 公告 API
2 extractor/ M2 GNE 中文新闻正文提取,输出 Article
3 dedup/ M3 三层去重:URL Hash → Content Hash → SimHash
4 llm/ M4 DeepSeek/Qwen 投资事件抽取,输出 ExtractedEvent
5 embedding/ M5 DashScope 远程 / 本地 BGE-M3 向量化
6 vectorstore/ M6 Qdrant 本地文件模式,语义检索
7 scheduler/ M7 APScheduler 定时任务 + pipeline 编排 + 日报生成
8 mcp_server/ M8 MCP 5 工具,供 Cherry Studio/Claude Code 调用
9 a_share_cli/ CLI 统一命令行入口(argparse),所有子命令
10 report_db/ M10 日报 MySQL 结构化入库(pymysql)
11 report_import/ M10 历史日报 HTML 解析与批量导入

辅助目录:

目录 用途
configs/ sources.yaml(新闻源)、watchlist.yaml(关注列表)、llm_models.yaml(LLM 场景配置)、loader.py
prompts/ 4 个 LLM Prompt 模板(event_extraction/company_analysis/industry_analysis/risk_analysis)
scripts/ 独立运行脚本(每个模块一个入口) + systemd 服务文件
tests/ pytest 测试(225+ passed)
api/ 占位(空包,预留给未来 API 服务)
app/ 空目录(预留)

2. 数据流全景

财经网站 / cninfo API
        │
        ▼
  ┌───────────┐
  │  M1 抓取   │ → data/raw/{source}/{YYYYMMDD}/*.html + index.jsonl
  └─────┬─────┘
        │
        ▼
  ┌───────────┐
  │  M2 提取   │ → data/processed/{source}/{YYYYMMDD}/{url_hash}.json  (Article)
  └─────┬─────┘
        │
        ▼
  ┌───────────┐
  │  M3 去重   │ → data/deduped/{YYYYMMDD}/uniques/{url_hash}.json
  └─────┬─────┘      data/dedup/fingerprints.sqlite3 (指纹库)
        │
        ▼
  ┌───────────┐
  │  M4 LLM   │ → data/events/{YYYYMMDD}/{url_hash}.json  (ExtractedEvent)
  └─────┬─────┘
        │
        ▼
  ┌───────────┐
  │ M5 向量化  │ → data/embeddings/{YYYYMMDD}/{url_hash}.json  (1024 维向量)
  └─────┬─────┘
        │
        ▼
  ┌───────────┐
  │ M6 入库    │ → data/qdrant_storage/ (Qdrant 本地文件)
  └─────┬─────┘
        │
   ┌────┴────┐
   ▼         ▼
┌──────┐  ┌──────┐
│ MCP  │  │ 日报  │
│ 检索  │  │ 生成  │
└──────┘  └──┬───┘
             │
             ▼
        ┌────────┐
        │ MySQL  │ → news_report / news_event (myquant 库)
        └────────┘

关键设计:

  • 增量处理:M2/M4/M5 产物存在即跳过(2026-08-12 新增),--force 全量重建
  • 断点续跑:pipeline --once --resume 从上次失败步骤继续,状态文件 data/pipeline/state.json
  • 去重多源记录:M3 指纹库 source_ids 列记录同一篇新闻的多个来源

3. 包职责与导出 API

3.1 crawler/ — M1 新闻抓取

文件 职责
engine.py Crawl4AI 异步引擎:httpx 静态直连(js_render=false) / Playwright(js_render=true)
config.py 加载 configs/sources.yaml
models.py SourceConfig / CrawlerConfig / CrawlResult / ArticleLink / CninfoItem / CrawlStage
storage.py 保存 raw HTML + index.jsonl,URL 去重
cninfo.py cninfo 公告/调研/互动易 API 抓取

公开 API:

from crawler import (
    crawl_all, crawl_source, extract_article_links,
    load_crawler_config,
    SourceConfig, CrawlerConfig, CrawlResult, CrawlStage, ArticleLink, CninfoItem,
)

抓取分流规则:

条件 引擎 特点
js_render=False 且无 wait_for httpx 直连 快,不触发反爬
js_render=True 或有 wait_for Playwright 支持 JS 渲染

新闻源:14 个(13 Web + 1 API:新闻联播),见 configs/sources.yaml

3.2 extractor/ — M2 正文提取

文件 职责
parser.py GNE 中文正文提取,模板文本过滤
models.py Article / ExtractError

公开 API:

from extractor import (
    extract_article, Article, ExtractError, MIN_CONTENT_LENGTH, SOURCE_NAME_MAP,
)

Article 模型(M2 输出,M3+ 输入的唯一格式):

字段 类型 说明
source_id str M1 源 ID
url str 文章 URL
url_hash str SHA1 前 16 位,主键
title str 标题
content str 清理后纯文本正文
author str|None 作者
publish_time datetime|None 标准化发布时间
word_count int 中文字符数
item_type str|None cninfo 类型:announcement/research/irm

3.3 dedup/ — M3 三层去重

文件 职责
deduper.py Deduper 主类:check / ingest / stats
hasher.py SimHash 64 位 + 汉明距离 + content_hash
store.py FingerprintStore,SQLite 持久化
models.py DedupLayer / DedupResult / DedupStats / Fingerprint

公开 API:

from dedup import (
    Deduper, article_to_fingerprint, FingerprintStore,
    DedupLayer, DedupResult, DedupStats, Fingerprint,
    simhash64, hamming, content_hash, normalize_content,
    DEFAULT_HAMMING_THRESHOLD, DEFAULT_TIME_WINDOW_DAYS,
)

三层去重逻辑:

层 算法 命中条件
L1 URL URL Hash (SHA1) 完全相同 URL
L2 Content 标准化正文 SHA1 正文完全一致
L3 SimHash 64 位 SimHash + 汉明距离 距离 ≤ 3 (默认)

3.4 llm/ — M4 投资事件抽取

文件 职责
client.py OpenAI 兼容客户端,指数退避重试,LLMConfig 加载
extractor.py Prompt 模板加载,JSON 解析,extract_event / extract_event_async
models.py EventExtraction / ExtractedEvent / Sentiment / LLMCallError / EVENT_TYPES

公开 API:

from llm import (
    extract_event, extract_event_async, parse_event_json,
    PromptTemplate, load_llm_config, LLMConfig,
    make_sync_client, make_async_client,
    EventExtraction, ExtractedEvent, Sentiment, LLMCallError,
    EVENT_TYPES, MAX_CONTENT_CHARS, MAX_IMPORTANCE, MIN_IMPORTANCE,
    SCENE_EVENT_EXTRACTION, SCENE_DAILY_REPORT, SCENE_STOCK_REPORT,
)

ExtractedEvent 模型(M4 落盘格式):

字段 类型 说明
source_id str 主源
url str 原文 URL
url_hash str SHA1 前 16 位
title str 标题
publish_time datetime|None 发布时间
sources list[str] 全部来源(去重合并)
event EventExtraction LLM 抽取结果
provider str deepseek / qwen
model str 模型名
attempts int LLM 实际调用次数(含重试)

EventExtraction(LLM 输出 JSON 结构):

字段 类型 说明
stock_codes list[str] 6 位代码,可带 .SH/.SZ/.BJ
company_names list[str] 公司中文简称
industries list[str] 行业(申万二级)
sentiment Sentiment positive/neutral/negative
importance int 1-5 重要程度
event_type str 23 种事件类型之一
summary str 一句话摘要(≤200 字)

23 种事件类型:业绩预告/业绩快报/财报披露/合作签约/投资并购/重大合同/产品发布/技术突破/监管处罚/诉讼仲裁/股东减持/股东增持/回购/分红/高管变动/资产重组/停牌复牌/ST警示/退市风险/宏观政策/行业政策/国际局势/其他

LLM 场景配置(configs/llm_models.yaml):

场景 用途 调用方
event_extraction 投资事件抽取(M4) llm/extractor.py
daily_report 日报 AI 摘要 scheduler/reporter.py
stock_report 个股 AI 要点 scheduler/stock_reporter.py
embedding 文本向量化(M5) embedding/factory.py

3.5 embedding/ — M5 向量化

文件 职责
base.py EmbeddingProvider / AsyncEmbeddingProvider ABC + compose_text
remote.py DashScope 远程 Embedding(OpenAI 兼容)
local.py 本地 BGE-M3(可选,需 uv sync --extra local-embedding)
factory.py resolve_provider_type / make_sync_provider / make_async_provider
models.py EmbeddingProviderType / EmbeddingResult / EmbeddingError

公开 API:

from embedding import (
    make_sync_provider, make_async_provider, resolve_provider_type,
    compose_text, EmbeddingProvider, AsyncEmbeddingProvider,
    EmbeddingResult, EmbeddingError, EmbeddingProviderType,
    DashScopeEmbeddingProvider, DashScopeAsyncEmbeddingProvider,
    DASHSCOPE_DEFAULT_DIM, DASHSCOPE_DEFAULT_MODEL, DASHSCOPE_BATCH_LIMIT,
    MAX_TEXT_CHARS,
)

向量维度:1024(DashScope text-embedding-v3 / 本地 BGE-M3)

文本组装:compose_text() 统一策略——标题 + 正文 + 事件摘要,最大 8000 字符

3.6 vectorstore/ — M6 Qdrant 知识库

文件 职责
client.py VectorStore 封装(init/upsert/query/count/info) + make_qdrant_client 工厂
models.py SearchFilter / SearchResult / CollectionInfo

公开 API:

from vectorstore import (
    VectorStore, make_qdrant_client,
    SearchFilter, SearchResult, CollectionInfo,
    DEFAULT_COLLECTION, DEFAULT_VECTOR_DIM,
)

Qdrant 模式:默认本地文件模式(data/qdrant_storage/),零依赖,ARM64 兼容

SearchFilter 支持:source_id / stock_codes / company_names / industries / sentiment / importance_min / event_types / publish_date_from / publish_date_to

3.7 scheduler/ — M7 定时任务与 Pipeline

文件 职责
pipeline.py run_pipeline / run_step,编排 M1→M6,断点续跑
reporter.py 日报生成(收集→AI 摘要→MySQL 入库)
stock_reporter.py 个股日报生成(关注列表)

公开 API:

from scheduler import (
    run_pipeline, run_step, PipelineResult, StepResult,
    STEP_COMMANDS, STEP_TIMEOUTS,
)

Pipeline 步骤:

步骤名 中文 默认超时 说明
crawler M1 新闻抓取 900s 13 源,含 Playwright
xwlb M1 新闻联播 60s 纯 HTTP API
extractor M2 正文提取 300s GNE 提取
dedup M3 新闻去重 120s 三层去重
llm M4 LLM 抽取 900s API 调用
embedding M5 向量化 300s DashScope
qdrant M6 Qdrant 入库 300s 本地文件写入
report 日报生成 30s 仅 07:00 执行
cninfo_crawl cninfo 公告抓取 900s 06:30 执行
cninfo_extract cninfo 正文提取 300s —
cninfo_pdf cninfo PDF 补充 — —

超时优先级:TIMEOUT_{NAME} 环境变量 > PIPELINE_STEP_TIMEOUT > 硬编码默认值 > 1800s

调度时间表:

时间 步骤
06:30 cninfo 公告管道
07:00 crawler→xwlb→extractor→dedup→llm→embedding→qdrant→日报
07:30 个股日报
12:00 crawler→xwlb→extractor→dedup→llm→embedding→qdrant
18:00 crawler→xwlb→extractor→dedup→llm→embedding→qdrant
22:00 crawler→xwlb→extractor→dedup→llm→embedding→qdrant

3.8 mcp_server/ — M8 MCP 服务

文件 职责
tools.py FastMCP 服务器,5 个 MCP 工具

5 个 MCP 工具:

工具 参数 说明
search_news query, top_k 通用语义检索
search_company_news query, company, top_k 按公司过滤
search_industry_news query, industry, top_k 按行业过滤
search_stock_events query, stock_code, top_k 按股票代码过滤
search_sentiment_trend query, sentiment, top_k 情绪趋势 + 统计

工具参数/返回格式/降级策略详见 docs/mcp_tools.md

启动方式:stdio 模式(Cherry Studio/Claude Code 自动管理进程)或 SSE 模式(调试:--sse 8765)

3.9 a_share_cli/ — 统一 CLI

文件 职责
main.py argparse 子命令路由,所有命令行操作

全部子命令:

命令 功能 关键参数
crawl M1 抓取 --source, --no-save
extract M2 提取 --date, --source
dedup M3 去重 --date, --reset
events M4 LLM 抽取 --date, --provider, --model, --limit, --concurrency
embed M5 向量化 --date, --provider, --model
ingest M6 入库 --date, --recreate
pipeline 全链路/守护 --once, --resume, --steps, --report, --cninfo-once
search 检索知识库 query, --top, --source, --sentiment, --stock, --industry, --min-importance
status 数据总览 无参数
report 生成日报 --date, --no-upload
report-import 历史日报导入 --dir, --date, --type, --force
cninfo 公告抓取 --no-save, --enrich-pdf, --pdf-limit
stock-report 个股日报 --no-upload
watchlist 关注列表管理 add/remove/list
discover 站点分析 url, --name, --extra, --add
add-entry 追加入口 source, urls...

3.10 report_db/ — M10 日报 DB 层

文件 职责
db.py 连接/事务/建表/保存/查询
models.py EventRow / ReportData (Pydantic)
schema.py DDL(CREATE TABLE)

公开 API:

from report_db import (
    connect, init_schema, save_report, load_db_config,
    exists_report, fetch_report,
    EventRow, ReportData,
)

3.11 report_import/ — M10 历史日报导入

文件 职责
parser.py BeautifulSoup 解析 finance/intl 历史 HTML
importer.py 批量导入,幂等,ImportStats 统计

公开 API:

from report_import import (
    import_history, ImportStats,
    parse_report, parse_finance_report, parse_intl_report, ReportParseError,
)

4. 核心数据模型

4.1 管道数据模型关系

CrawlResult (M1) ──→ Article (M2) ──→ DedupResult (M3)
                                            │
                                            ▼
                                     ExtractedEvent (M4)
                                     │
                                     ▼
                                     EmbeddingResult (M5)
                                     │
                                     ▼
                                     SearchResult (M6)

4.2 日报数据模型

ReportData (Pydantic)
├── report_date: date
├── report_type: str (finance | intl)
├── file_name: str
├── generated_at: datetime
├── ai_summary: str | None
├── stats: dict (JSON)
└── events: list[EventRow]
    ├── section: str (xwlb | news | cninfo | intl)
    ├── rank: int
    ├── importance: int | None
    ├── event_type: str | None
    ├── title: str
    ├── summary: str | None
    ├── sentiment: str | None
    └── source: str | None

4.3 全部 Pydantic 模型清单

包 模型 用途
crawler SourceConfig 单个新闻源配置
crawler CrawlerConfig 全局抓取配置
crawler CrawlerSettings 抓取设置(并发/重试/UA)
crawler CrawlResult 单次抓取结果
crawler ArticleLink 列表页发现的链接
crawler CninfoItem 公告/调研/互动易条目
extractor Article 标准化文章(M2 输出)
dedup Fingerprint 指纹记录(SQLite)
dedup DedupResult 判重结果
dedup DedupStats 指纹库统计
llm EventExtraction LLM 输出 JSON 结构
llm ExtractedEvent M4 落盘格式(文章+事件)
embedding EmbeddingResult 向量化结果
vectorstore SearchFilter 检索过滤条件
vectorstore SearchResult 单条检索结果
vectorstore CollectionInfo Qdrant Collection 信息
scheduler StepResult 单步执行结果
scheduler PipelineResult 全链路执行结果
report_db EventRow 日报事件行
report_db ReportData 完整日报数据

5. 配置体系

5.1 配置文件

文件 格式 用途 热更新
configs/sources.yaml YAML 14 个新闻源配置 每次抓取重读
configs/llm_models.yaml YAML 4 个 LLM 场景配置 每次调用重读(mtime 缓存失效)
configs/watchlist.yaml YAML cninfo 公告关注列表 每次操作重读
.env dotenv API Key + 调度/超时/DB 配置 ~2s 内自动生效(常驻进程无需重启)

热加载实现见 configs/runtime_env.py:常驻进程(调度器 / MCP server)启动后 会以 2s 周期比对 .env 的 (mtime, size),变化即写入 os.environ; configs/loader.py 对 YAML 做同样的 mtime 失效。 进程环境里显式设置且与文件不同的变量优先(如 LLM_PROVIDER=qwen ...), .env 中删除的键也会同步从环境中移除。

5.2 配置优先级(LLM 场景)

CLI 显式参数 (--provider / --model)
    ↓
configs/llm_models.yaml scenes.<场景>.xxx
    ↓
.env 环境变量 (LLM_PROVIDER / DEEPSEEK_MODEL 等)
    ↓
代码内置默认值 (仅超时/温度等参数)

模型名无内置兜底:缺失即报错,绝不静默使用错误模型。

5.3 关键环境变量

变量 用途 生产值示例
DASHSCOPE_API_KEY 百炼 LLM + Embedding sk-xxx
DEEPSEEK_API_KEY DeepSeek LLM sk-xxx
SCHEDULE_TIMES 调度时间 07:00,12:00,18:00,22:00
PIPELINE_STEP_TIMEOUT 全局超时 1800
NEWS_DB_HOST 日报 DB 主机 192.168.1.10(pi5) / 127.0.0.1(Mac)
NEWS_DB_PORT 日报 DB 端口 13306

6. 产物目录结构

data/
├── raw/                           # M1 原始抓取
│   ├── cls/{YYYYMMDD}/            #   按源分目录
│   │   ├── index.jsonl            #     文章列表(CrawlResult 摘要)
│   │   └── *.html                 #     原始 HTML
│   ├── eastmoney/{YYYYMMDD}/
│   ├── ...
│   └── cninfo/{YYYYMMDD}/         #   公告独立目录
│
├── processed/                     # M2 正文提取
│   ├── cls/{YYYYMMDD}/
│   │   └── {url_hash}.json        #   Article
│   └── ...
│
├── dedup/                         # M3 指纹库(跨日)
│   └── fingerprints.sqlite3
│
├── deduped/                       # M3 去重产物
│   └── {YYYYMMDD}/
│       ├── uniques/
│       │   └── {url_hash}.json    #   唯一文章
│       ├── duplicates.jsonl       #   重复记录
│       └── sources.json           #   多源映射
│
├── events/                        # M4 事件抽取
│   └── {YYYYMMDD}/
│       └── {url_hash}.json        #   ExtractedEvent
│
├── embeddings/                    # M5 向量化
│   └── {YYYYMMDD}/
│       └── {url_hash}.json        #   1024 维向量 + 事件
│
├── qdrant_storage/                # M6 Qdrant 本地文件
│
├── pipeline/                      # 断点续跑状态
│   └── state.json
│
└── reports_history/               # 历史日报 HTML(可选)

7. 运行环境

项目 值
Python 3.11
包管理器 uv
虚拟环境 .venv/
生产服务器 树莓派 5 (ARM64), pi@192.168.1.160
数据库 MySQL/MariaDB (myquant 库), 通过 pi@192.168.1.10 autossh 隧道
系统服务 systemd a-share-research + a-share-db-tunnel
测试 pytest (asyncio_mode=auto, integration 标记默认跳过)
静态检查 ruff + mypy
LLM 模型 deepseek-v4-flash(锁定,不可修改)

—— 架构文档结束 ——