docs: 文档重构 — 清理 AI agent 残留,整合 docs/ 目录结构
- 删除 11 个残留文件: continuation.md, init_plan.md, reasonix.toml, djapi/continuation.md, djapi/.serena/, djapi/.claude/, djapi/.mcp.json, .claude/skills/, docs/usage.html, docs/db_schema.md, docs/report_db_design.md - 7 个 CLAUDE-*.md 移入 docs/ 并重命名去 CLAUDE- 前缀 - 新增 4 个文档: architecture.md, development.md, api.md, deployment.md - 重写 usage.md, README.md - 修复所有过时引用和交叉链接
This commit is contained in:
@@ -1,26 +0,0 @@
|
||||
# batch-sync skill
|
||||
|
||||
批量预热股票数据到 DB 缓存。
|
||||
|
||||
## 触发
|
||||
|
||||
用户说:预热缓存 / 同步数据 / warmup / batch sync / 补齐数据 / 全量同步
|
||||
|
||||
## 执行
|
||||
|
||||
```bash
|
||||
cd finance && python cli/agent_cli.py warmup 50
|
||||
```
|
||||
|
||||
## 说明
|
||||
|
||||
- 每批 50 只股票,依次执行 `dm.sync_daily()`
|
||||
- 已缓存 + 最新的 → 0 条跳过(增量)
|
||||
- 未缓存 → Tushare 优先 → AkShare fallback
|
||||
- 范围由 `.env` 中 `SENTIMENT_SCOPE_TYPE` + `SENTIMENT_SCOPE_INDEXES` 决定(默认沪深300+中证500)
|
||||
- 可多次执行直到覆盖率 100%
|
||||
|
||||
## 参数
|
||||
|
||||
`python cli/agent_cli.py warmup [N]`
|
||||
- N: 每批股票数,默认 50。网络稳定时可调大到 100
|
||||
@@ -7,225 +7,123 @@
|
||||
```
|
||||
cc-cursor/
|
||||
├── finance/ # 核心量化引擎
|
||||
│ ├── config/ # 全局配置(MariaDB / AkShare)
|
||||
│ ├── database/ # ORM 模型 + DAO(mac_ 前缀表)
|
||||
│ ├── config/ # 全局配置
|
||||
│ ├── database/ # ORM 模型 + DAO
|
||||
│ ├── data/ # DataManager 统一数据层
|
||||
│ ├── factors/ # 因子引擎(34 因子 / 12 分类,含情绪因子)
|
||||
│ ├── backtest/ # 回测引擎(VectorBT + 5 策略 + 截面回测)
|
||||
│ ├── optimizer/ # Optuna 参数优化 + Walk-Forward
|
||||
│ ├── factors/ # 因子引擎(34 因子 / 12 分类)
|
||||
│ ├── backtest/ # 回测引擎(VectorBT + 5 策略)
|
||||
│ ├── optimizer/ # Optuna 参数优化
|
||||
│ ├── models/ # LightGBM / CatBoost ML 模型
|
||||
│ ├── agents/ # Agent 系统(4 Agent + 编排器 + CLI)
|
||||
│ ├── agents/ # Agent 系统(4 Agent + 编排器)
|
||||
│ ├── cli/ # 命令行 & 验证脚本
|
||||
│ ├── reports/ # 自动日报输出目录
|
||||
│ └── .env # 环境变量配置(API Key / 分析范围)
|
||||
├── djapi/ # Django API 后端(A 股数据 + 新闻联播 + 日报查询)
|
||||
├── mcp-servers/ # MCP Server(Serena,本机工具,git 忽略)
|
||||
├── shared/ # 共享工具(SSH 隧道脚本)
|
||||
├── docs/ # 文档 & 使用指南
|
||||
└── .claude/ # Claude Code 配置
|
||||
│ └── reports/ # 日报输出目录
|
||||
├── djapi/ # Django API 后端
|
||||
├── shared/script/ # SSH 隧道脚本
|
||||
└── docs/ # 项目文档
|
||||
```
|
||||
|
||||
## 数据流
|
||||
|
||||
```
|
||||
Agent 编排层
|
||||
├── ResearchAgent ── 因子发现(IC/IC_IR 评估)
|
||||
├── SelectionAgent ─ 多因子打分 + ML 预测
|
||||
├── RiskAgent ────── 仓位控制 + 风险预警
|
||||
└── ReportAgent ──── 自动日报生成
|
||||
|
||||
基础引擎层
|
||||
DataManager ──→ FactorEngine ──→ BaseStrategy ──→ VectorBTEngine ──→ BacktestReport
|
||||
│ │ │
|
||||
│ FeatureEngine OptunaEngine
|
||||
│ │ │
|
||||
└──────→ LightGBM/CatBoost ←────────┘
|
||||
|
||||
情绪增强层
|
||||
NewsSource(AkShare/DB/MCP) ──→ QwenClient ──→ SentimentFactor ──→ FactorEngine
|
||||
Data → Factor → Model → Strategy → Backtest → Report
|
||||
```
|
||||
|
||||
全部通过 Service 层中转:策略不直连 AkShare,模型不直连数据库,Agent 不重建引擎。
|
||||
|
||||
---
|
||||
|
||||
## 开发进度
|
||||
|
||||
| Sprint | 模块 | 关键成果 | 状态 |
|
||||
|--------|------|----------|------|
|
||||
| Sprint 0 | 基础设施 | DataManager + MariaDB 3 表 | ✅ |
|
||||
| Sprint 1 | 因子引擎 | 34 因子 / 12 分类 | ✅ |
|
||||
| Sprint 2 | 回测引擎 | VectorBT + 5 策略 + 截面回测 | ✅ |
|
||||
| Sprint 3 | 参数优化 | Optuna + Walk-Forward | ✅ |
|
||||
| Sprint 4 | ML 模型 | LightGBM + CatBoost + 特征工程 | ✅ |
|
||||
| Sprint 5 | 情绪因子 | Qwen + 三源新闻聚合 + 日期对齐 | ✅ |
|
||||
| Sprint 6 | Agent 系统 | 4 Agent + 编排器 + CLI + 自动日报 | ✅ |
|
||||
| Sprint 7 | djapi API | 日报查询 ×2(news/reports + news/events) | ✅ |
|
||||
|
||||
**全部 8 个 Sprint 已完成。**
|
||||
|
||||
---
|
||||
|
||||
## 功能模块
|
||||
|
||||
### 数据层 `finance/data/`
|
||||
|
||||
```python
|
||||
from data.data_manager import DataManager
|
||||
dm = DataManager(); dm.init_db()
|
||||
stocks = dm.get_stock_list() # → 5,524 只
|
||||
daily = dm.get_daily("000001.SZ") # → 日线
|
||||
fina = dm.get_financial("000001.SZ") # → 财务数据
|
||||
dm.sync_daily("000001.SZ") # → 增量同步
|
||||
```
|
||||
|
||||
### 因子引擎 `finance/factors/`
|
||||
|
||||
```python
|
||||
from factors.registry import get_factor, list_factors
|
||||
from factors.engine import FactorEngine
|
||||
|
||||
engine = FactorEngine(dm)
|
||||
factors = [get_factor("momentum_20"), get_factor("rsi_14")]
|
||||
factor_df = engine.compute("000001.SZ", factors)
|
||||
# → 34 个注册因子,12 个分类(动量/RSI/MACD/量价/布林/ATR/均线/波动率/换手率/振幅/基本面/情绪)
|
||||
```
|
||||
|
||||
### 回测引擎 `finance/backtest/`
|
||||
|
||||
```python
|
||||
from backtest.vectorbt.engine import VectorBTEngine
|
||||
from backtest.strategies.rsi_mean_revert import RSIMeanRevertStrategy
|
||||
|
||||
engine_bt = VectorBTEngine(initial_capital=100_000, commission=0.0003)
|
||||
report = engine_bt.run(RSIMeanRevertStrategy(oversold=30, overbought=70), price_df, factor_df)
|
||||
# → 收益=29.4% 年化=4.3% 回撤=-19.1% 夏普=0.37 胜率=77.1%
|
||||
```
|
||||
|
||||
5 个内置策略 + 自定义策略接口 + 截面回测 + BacktestReport 标准化报告。
|
||||
|
||||
### 参数优化 `finance/optimizer/`
|
||||
|
||||
```python
|
||||
from optimizer.engine import OptunaEngine
|
||||
from optimizer.space import rsi_revert_space
|
||||
|
||||
result = OptunaEngine(engine_bt).optimize(
|
||||
RSIMeanRevertStrategy, rsi_revert_space, price_df, factor_df,
|
||||
metric="sharpe", n_trials=200,
|
||||
)
|
||||
# → 最优参数: oversold=13, overbought=66
|
||||
# → 夏普: 0.37→0.60 (+62%), 回撤: -19.1%→-1.8% (10倍改善)
|
||||
```
|
||||
|
||||
7 种优化目标 + 4 个预置搜索空间 + Walk-Forward 滚动验证 + 快捷函数。
|
||||
|
||||
### ML 模型 `finance/models/`
|
||||
|
||||
```python
|
||||
from models.features import FeatureEngine
|
||||
from models.lightgbm.model import LightGBMModel
|
||||
|
||||
fe = FeatureEngine(lookahead=5)
|
||||
X, y = fe.build(factor_df, price_df, fit=True)
|
||||
model = LightGBMModel(params={"n_estimators": 200}).fit(X_train, y_train)
|
||||
pred = model.predict(X_test) # → IC 评估 + 特征重要性 + 交叉验证 + ML 策略回测
|
||||
```
|
||||
|
||||
Winsorize → 缺失填充 → RobustScaler → LightGBM/CatBoost 训练 → MLBenchmark 对比。
|
||||
|
||||
### 情绪因子 `finance/factors/sentiment/`
|
||||
|
||||
```python
|
||||
from factors.sentiment.sentiment_engine import SentimentEngine
|
||||
|
||||
sent = SentimentEngine(dm)
|
||||
sent_df = sent.compute("000001.SZ", max_news=20)
|
||||
# → news_sent_5, news_conf_5, sent_delta_5
|
||||
```
|
||||
|
||||
三数据源聚合(AkShare 个股新闻 + 新闻联播 DB + MCP trendradar-news)、日期对齐(非交易日→最近交易日)、xwlb 偏移(昨日新闻→今日使用)、DashScope + Ollama 双后端。
|
||||
|
||||
### Agent 系统 `finance/agents/`
|
||||
|
||||
```bash
|
||||
python finance/cli/agent_cli.py daily # 完整每日流程
|
||||
python finance/cli/agent_cli.py picks 15 # 选股 Top 15
|
||||
python finance/cli/agent_cli.py risk # 风险评估
|
||||
python finance/cli/agent_cli.py research # 因子研究
|
||||
python finance/cli/agent_cli.py report # 生成日报
|
||||
```
|
||||
|
||||
4 个 Agent(Research/Selection/Risk/Report)+ 编排器 + 自动日报(reports/daily_YYYYMMDD.md)。
|
||||
|
||||
---
|
||||
|
||||
## 快速开始
|
||||
|
||||
```bash
|
||||
# SSH 隧道
|
||||
# 环境 & SSH 隧道
|
||||
conda activate quant
|
||||
bash shared/script/autossh.sh
|
||||
|
||||
# Python 环境
|
||||
conda activate quant # Python 3.11.13
|
||||
|
||||
# 每日 Agent 运行
|
||||
python finance/cli/agent_cli.py daily
|
||||
```
|
||||
|
||||
### 验证脚本
|
||||
## CLI 命令
|
||||
|
||||
```bash
|
||||
python finance/cli/demo_data_manager.py # Sprint 0 — DataManager
|
||||
python finance/cli/demo_factor_engine.py # Sprint 1 — 因子引擎
|
||||
python finance/cli/demo_backtest.py # Sprint 2 — 回测引擎
|
||||
python finance/cli/demo_optimizer.py # Sprint 3 — 参数优化
|
||||
python finance/cli/demo_ml.py # Sprint 4 — ML 模型
|
||||
python finance/cli/demo_sentiment.py # Sprint 5 — 情绪因子
|
||||
python finance/cli/demo_sentiment_detail.py # Sprint 5 — 情绪因子(单股详情)
|
||||
python finance/cli/agent_cli.py daily # 5 步完整流程
|
||||
python finance/cli/agent_cli.py picks 15 # 选股 Top 15
|
||||
python finance/cli/agent_cli.py risk # 风险评估
|
||||
python finance/cli/agent_cli.py research # 因子研究
|
||||
python finance/cli/agent_cli.py report 20260603 # 生成日报
|
||||
python finance/cli/agent_cli.py warmup 50 # 首次预热缓存
|
||||
```
|
||||
|
||||
---
|
||||
## 验证脚本
|
||||
|
||||
```bash
|
||||
python finance/cli/demo_data_manager.py --ts_code 600519.SH
|
||||
python finance/cli/demo_factor_engine.py --ts_code 300750.SZ
|
||||
python finance/cli/demo_backtest.py --ts_code 000001.SZ
|
||||
python finance/cli/demo_optimizer.py --ts_code 000001.SZ --trials 100
|
||||
python finance/cli/demo_ml.py --ts_code 000001.SZ --lookahead 5
|
||||
python finance/cli/demo_sentiment.py --ts_code 600519.SH
|
||||
python finance/cli/demo_sentiment_detail.py --ts_code 600519.SH --date 20260603
|
||||
```
|
||||
|
||||
## 技术栈
|
||||
|
||||
| 组件 | 技术 | 版本 | 状态 |
|
||||
|------|------|------|------|
|
||||
| 数据获取 | AkShare | 1.18.64 | ✅ |
|
||||
| 数据库 | MariaDB (SSH 隧道) | — | ✅ |
|
||||
| 因子/特征 | pandas / numpy / sklearn | 2.3 / 2.0 / 1.9 | ✅ |
|
||||
| 回测引擎 | VectorBT | 1.0 | ✅ |
|
||||
| 参数优化 | Optuna | 4.9 | ✅ |
|
||||
| ML 模型 | LightGBM / CatBoost | 4.6 / 1.2 | ✅ |
|
||||
| NLP 情绪 | Qwen (DashScope / Ollama) | turbo / 2.5 | ✅ |
|
||||
| Agent 框架 | 自研编排器 | — | ✅ |
|
||||
| API 后端 | Django + uWSGI | 5.2 | 已有 |
|
||||
| 代码分析 | Serena MCP | — | 本机工具(不随仓库分发) |
|
||||
| 组件 | 技术 |
|
||||
|------|------|
|
||||
| 数据获取 | AkShare + Tushare (双源) |
|
||||
| 数据库 | MariaDB (SSH 隧道) |
|
||||
| 因子/特征 | pandas / numpy / sklearn |
|
||||
| 回测引擎 | VectorBT 1.0 |
|
||||
| 参数优化 | Optuna 4.9 |
|
||||
| ML 模型 | LightGBM 4.6 + CatBoost 1.2 |
|
||||
| NLP 情绪 | Qwen (DashScope / Ollama) |
|
||||
| Agent 编排 | 自研编排器 |
|
||||
| API 后端 | Django 5.2 + uWSGI |
|
||||
|
||||
---
|
||||
## 开发进度
|
||||
|
||||
## 设计原则
|
||||
| Sprint | 模块 | 状态 |
|
||||
|--------|------|------|
|
||||
| Sprint 0 | 基础设施(DataManager + MariaDB) | ✅ |
|
||||
| Sprint 1 | 因子引擎(34 因子 / 12 分类) | ✅ |
|
||||
| Sprint 2 | VectorBT 回测(5 策略 + 截面) | ✅ |
|
||||
| Sprint 3 | Optuna 优化(+ Walk-Forward) | ✅ |
|
||||
| Sprint 4 | ML 模型(LightGBM + CatBoost) | ✅ |
|
||||
| Sprint 5 | Qwen 情绪因子(三源新闻) | ✅ |
|
||||
| Sprint 6 | Agent 系统(4 Agent + CLI) | ✅ |
|
||||
| Sprint 7 | djapi API(日报查询 ×2) | ✅ |
|
||||
|
||||
- **模块隔离**:各引擎通过统一接口交互,可替换实现(VectorBT → Backtrader)
|
||||
- **接口标准化**:因子 `calculate(df)→Series` / 策略 `generate_signals(df)→Series` / 模型 `fit/predict/save/load` / 优化 `optimize()→Result`
|
||||
- **数据层统一**:策略/模型不直连数据源,全部通过 DataManager
|
||||
- **Agent 不重建轮子**:Agent 通过依赖注入复用已有引擎,编排而非重建
|
||||
- **防前视偏差**:时间序列交叉验证、expanding window 统计量
|
||||
- **渐进演进**:全链路 8 个 Sprint 平滑推进,无推倒重写
|
||||
**全部 8 个 Sprint 已完成。**
|
||||
|
||||
## 文档
|
||||
|
||||
- [使用指南](./docs/usage.md) — 详细使用说明(12 章节,含代码示例)
|
||||
- [使用指南 (HTML)](./docs/usage.html) — 网页版使用指南
|
||||
- [新闻日报 API](./docs/news_report_api.md) — djapi 日报查询接口使用手册
|
||||
| 文档 | 内容 |
|
||||
|------|------|
|
||||
| [使用指南](docs/usage.md) | 各模块使用方法和代码示例 |
|
||||
| [架构说明](docs/architecture.md) | 项目架构、数据流、设计原则 |
|
||||
| [开发指南](docs/development.md) | 环境搭建、开发约定、模块说明 |
|
||||
| [部署说明](docs/deployment.md) | 本地环境、服务器、uWSGI、rsync 部署 |
|
||||
| [因子与表结构速查](docs/reference.md) | 34 因子注册表、DB 表结构 |
|
||||
| [数据层详解](docs/data-layer.md) | DataManager、数据库、缓存策略 |
|
||||
| [因子引擎详解](docs/factors.md) | 因子计算、情绪引擎、新闻源 |
|
||||
| [回测引擎详解](docs/backtest.md) | VectorBT、策略、信号工具、Optuna |
|
||||
| [ML 模型详解](docs/ml-models.md) | 特征工程、LightGBM/CatBoost |
|
||||
| [Agent 系统详解](docs/agents.md) | Agent 架构、CLI、日报 |
|
||||
| [DJAPI 接口](docs/api.md) | Django API 端点参考 |
|
||||
| [日报查询 API](docs/news_report_api.md) | news/reports + news/events 接口 |
|
||||
| [日报数据库](docs/db_schema_v1.1.md) | news_report / news_event 表结构 |
|
||||
|
||||
## 子项目
|
||||
|
||||
- [djapi](./djapi/README.md) — Django API 后端:A 股数据 API(16 端点)+ 新闻联播处理 + 日报查询(news/reports、news/events)
|
||||
- [djapi](djapi/README.md) — Django API 后端:A 股数据 API(16 端点)+ 新闻联播处理 + 日报查询
|
||||
|
||||
## 设计原则
|
||||
|
||||
- **模块隔离**:各引擎通过统一接口交互,可替换实现
|
||||
- **接口标准化**:因子 `calculate(df)→Series` / 策略 `generate_signals(df)→Series` / 模型 `fit/predict/save/load`
|
||||
- **数据层统一**:策略/模型不直连数据源,全部通过 DataManager
|
||||
- **Agent 不重建轮子**:Agent 通过依赖注入复用已有引擎
|
||||
- **防前视偏差**:时间序列交叉验证、expanding window 统计量
|
||||
|
||||
## 数据库连接
|
||||
|
||||
```bash
|
||||
bash shared/script/autossh.sh
|
||||
# host: 127.0.0.1:13306 user: myquant database: myquant table_prefix: mac_
|
||||
```
|
||||
```
|
||||
@@ -1,83 +0,0 @@
|
||||
# continuation.md — cc-cursor 项目状态
|
||||
|
||||
生成时间:2026-06-07(全部 Sprint 完成 + 生产加固 + djapi 数据源归一化 + Git 初始化)
|
||||
|
||||
---
|
||||
|
||||
## Git 状态
|
||||
|
||||
- 仓库:https://github.com/Simon2046/myquant
|
||||
- 分支:`main`
|
||||
- commit:`271a934` — Initial commit: cc-cursor 全链路量化研究平台
|
||||
- 文件:293 个文件,59,598 行
|
||||
- 已排除:`.env`、`mcp-servers/serena`、`__pycache__`、`.parquet`、`.db`
|
||||
|
||||
---
|
||||
|
||||
## 全部 Sprint 完成 ✅
|
||||
|
||||
| Sprint | 模块 | 状态 |
|
||||
|--------|------|------|
|
||||
| 0 | 基础设施(DataManager + MariaDB) | ✅ |
|
||||
| 1 | 因子引擎(34 因子 / 12 分类) | ✅ |
|
||||
| 2 | VectorBT 回测(5 策略 + 截面 + BacktestReport) | ✅ |
|
||||
| 3 | Optuna 优化(+ Walk-Forward) | ✅ |
|
||||
| 4 | ML 模型(LightGBM + CatBoost + MLStrategy) | ✅ |
|
||||
| 5 | Qwen 情绪因子(三源新闻 + 日期对齐) | ✅ |
|
||||
| 6 | Agent 系统(4 Agent + CLI + 日报 .md/.html) | ✅ |
|
||||
|
||||
---
|
||||
|
||||
## 生产稳定性加固(14 项)
|
||||
|
||||
| # | 项 | 文件 |
|
||||
|---|-----|------|
|
||||
| 1 | Tushare 双数据源(优先) | `finance/data/data_manager.py` |
|
||||
| 2 | 指数 vs 个股自动路由 | `finance/data/sources/akshare_source.py` |
|
||||
| 3 | SSH 自动恢复(多次重连 + pool_pre_ping) | `finance/database/connection.py` |
|
||||
| 4 | save_daily 先删后插(防主键冲突) | `finance/database/dao.py` |
|
||||
| 5 | load_dotenv 绝对路径 + 模块加固 | `finance/config/settings.py` + 3 文件 |
|
||||
| 6 | 日报 5d/20d 修复(idx=-1→pos=len-1) | `finance/agents/report_agent.py` |
|
||||
| 7 | RiskAgent 改用上证指数 | `finance/agents/risk_agent.py` |
|
||||
| 8 | 日报增加"昨日对比" + 数据截止 | `finance/agents/report_agent.py` |
|
||||
| 9 | mac_report 表 utf8mb4 + DATE + DATETIME | `finance/database/models.py` |
|
||||
| 10 | 日报自动存入 DB + emoji 兼容 | `finance/reports/storage.py` + 8 CLI |
|
||||
| 11 | CLAUDE-*.md 文档化 9 条已知 Bug | `CLAUDE-data.md` + `CLAUDE-agents.md` |
|
||||
| 12 | demo 脚本全参数化 | `finance/cli/demo_*.py` |
|
||||
| 13 | **djapi 数据源归一化(10→1 入口)** | `djapi/api/stock/data_source.py` |
|
||||
| 14 | **indexDatas API 参数修正 + 容错** | `djapi/api/views.py` + `getIndexs.py` |
|
||||
|
||||
---
|
||||
|
||||
## djapi 数据源归一化
|
||||
|
||||
- 新增 `djapi/api/stock/data_source.py` — 统一入口
|
||||
- `get_tushare_pro()` — 全局单例(线程安全)
|
||||
- `get_daily()` — 双源 fallback (Tushare→AkShare)
|
||||
- `get_mysql_db()` — MySQL 全局单例
|
||||
- Token 兼容 `TUSHARE_TS_TOKEN` / `TUSHARE_TOKEN`
|
||||
- 10 个模块迁移完成
|
||||
- `getDivData_AK.py` 标记废弃
|
||||
- `getIndexs.py` 修复:`index_dailybasic` 失败不阻塞,异常 raise 而非静默返回空
|
||||
- `views.py` 修正:`indexDatas` 参数 `index_name` → `tscode`,描述从"股票代码"→"指数代码",新增 `_PARAM_INDEX_CODE`
|
||||
- 已部署到 `api.doorcome.cn` ✅
|
||||
|
||||
---
|
||||
|
||||
## CLI 命令
|
||||
|
||||
```bash
|
||||
agent_cli.py daily / picks / risk / research / report / warmup
|
||||
demo_*.py(全部支持 --ts_code --date 等参数)
|
||||
```
|
||||
|
||||
## 文档
|
||||
|
||||
- `CLAUDE.md` — 入口 + 路由 + 多步任务规则
|
||||
- `CLAUDE-data.md` — 数据层 + 5 条已知 Bug
|
||||
- `CLAUDE-factors.md` — 因子引擎
|
||||
- `CLAUDE-backtest.md` — 回测 + 优化
|
||||
- `CLAUDE-ml.md` — ML 模型
|
||||
- `CLAUDE-agents.md` — Agent + CLI + 4 条已知 Bug
|
||||
- `CLAUDE-reference.md` — 因子/表结构速查
|
||||
- `docs/usage.md` + `docs/usage.html` — 使用指南
|
||||
+1
-1
@@ -13,7 +13,7 @@ MYSQL_PASSWORD=your-mysql-password
|
||||
MYSQL_DATABASE=myquant
|
||||
|
||||
# 日报结构化入库 (news_report / news_event) 只读查询
|
||||
# 与 report_db_design.md §7 一致;密码必填,缺失时接口直接报错
|
||||
# 与 docs/db_schema_v1.1.md 一致;密码必填,缺失时接口直接报错
|
||||
NEWS_DB_HOST=127.0.0.1
|
||||
NEWS_DB_PORT=3306
|
||||
NEWS_DB_USER=myquant
|
||||
|
||||
@@ -1,16 +0,0 @@
|
||||
{
|
||||
"mcpServers": {
|
||||
"serena-djapi": {
|
||||
"command": "uv",
|
||||
"args": [
|
||||
"run",
|
||||
"--directory",
|
||||
"/Users/summer/Downloads/cc-cursor/mcp-servers/serena",
|
||||
"serena",
|
||||
"start-mcp-server",
|
||||
"--project",
|
||||
"/Users/summer/Downloads/cc-cursor/djapi"
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,2 +0,0 @@
|
||||
/cache
|
||||
/project.local.yml
|
||||
@@ -1,23 +0,0 @@
|
||||
# Code Style & Conventions
|
||||
|
||||
## Python
|
||||
- Django app: all business logic in `api/stock/`, not in views
|
||||
- views.py is thin forwarding layer: extract params -> call function -> return Response
|
||||
- Double import pattern for standalone scripts: try relative import first, fall back to absolute
|
||||
- Use `viewFunc_tsCodeAndDate()` wrapper for ts_code + date_range endpoints
|
||||
- Use `viewFunc_singleParam()` wrapper for single-param endpoints
|
||||
- DRF `@api_view(['GET'])` + `@extend_schema` on all views
|
||||
- DRF `Response` (not `JsonResponse`) — no `safe=False` parameter
|
||||
- Configuration split: config.py (token) / strategy_config.py / scan_config.py
|
||||
|
||||
## Constraints
|
||||
- `api/video/` is protected — do NOT modify unless user explicitly asks
|
||||
- No python-dotenv dependency — use stdlib env loaders only
|
||||
- Backward compatibility: keep re-exports when splitting modules
|
||||
- Server `.env` file manages all secrets; uwsgi.ini only has DJANGO_SETTINGS_MODULE
|
||||
|
||||
## Secrets
|
||||
- All API keys/tokens/passwords via os.getenv()
|
||||
- Local: .env file (not committed)
|
||||
- Server: /home/simon/myquant/djapi/.env
|
||||
- Django loads via djapi/env_loader.py, video loads via api/video/env.py
|
||||
@@ -1,32 +0,0 @@
|
||||
# Project Overview
|
||||
|
||||
djapi is a Django 5.2 project providing financial data APIs for A-share stocks and CCTV news broadcast video processing.
|
||||
|
||||
## Tech Stack
|
||||
- Python 3.10, Django 5.2, uWSGI, nginx
|
||||
- Tushare (stock data), akshare (alternative stock data)
|
||||
- DRF + drf-spectacular (API documentation)
|
||||
- MySQL (business data), SQLite (Django admin only)
|
||||
- yt-dlp + ffmpeg + pydub (video/audio processing)
|
||||
- DashScope (ASR), DeepSeek API (AI text processing)
|
||||
|
||||
## Architecture
|
||||
- Single Django app: `api`
|
||||
- `api/stock/` — stock data module (Tushare/akshare -> pandas -> JsonResponse/DRF Response)
|
||||
- `api/video/` — independent video processing pipeline (download -> audio -> ASR -> AI split -> MySQL)
|
||||
- views.py is thin: extracts params, calls stock functions, returns Response
|
||||
|
||||
## Key Files
|
||||
- `api/views.py` — all ~15 API views, using @api_view + @extend_schema
|
||||
- `api/stock/stock_utils.py` — shared utilities: tscodeCheck, viewFunc_tsCodeAndDate, viewFunc_singleParam
|
||||
- `api/stock/config.py` — Tushare token + re-exports from strategy_config, scan_config
|
||||
- `api/serializers.py` — 13 DRF Serializer classes
|
||||
- `djapi/env_loader.py` — .env file loader (stdlib, no python-dotenv)
|
||||
- `api/video/env.py` — standalone .env loader for video module
|
||||
- `api/utils/mysql_handler.py` — shared MySQLDB class
|
||||
|
||||
## Deployment
|
||||
- Server: simon@doorcome.cn, path: /home/simon/myquant/djapi/
|
||||
- Virtual env: /opt/miniconda/envs/django/
|
||||
- uWSGI on port 5004, nginx reverse proxy
|
||||
- Domains: api.doorcome.cn, echart.doorcome.cn
|
||||
@@ -1,44 +0,0 @@
|
||||
# Suggested Commands
|
||||
|
||||
## Development
|
||||
```bash
|
||||
python manage.py runserver 0.0.0.0:8000 # dev server
|
||||
python manage.py check --deploy # check config
|
||||
python manage.py test api # run tests
|
||||
```
|
||||
|
||||
## uWSGI
|
||||
```bash
|
||||
uwsgi --ini uwsgi.ini # start
|
||||
uwsgi --reload uwsgi.pid # hot reload
|
||||
uwsgi --stop uwsgi.pid # stop
|
||||
# On server:
|
||||
/opt/miniconda/envs/django/bin/uwsgi --ini /home/simon/myquant/djapi/uwsgi.ini
|
||||
kill $(lsof -ti:5004) # force stop
|
||||
```
|
||||
|
||||
## Deploy
|
||||
```bash
|
||||
# Full sync (exclude production data)
|
||||
rsync -avz --delete \
|
||||
--exclude='.env' --exclude='db.sqlite3' \
|
||||
--exclude='*.log' --exclude='uwsgi.pid' \
|
||||
--exclude='__pycache__/' --exclude='*.pyc' \
|
||||
--exclude='xwlb_video/' --exclude='audio_processing/' \
|
||||
/Users/summer/Downloads/cc-cursor/djapi/ \
|
||||
simon@doorcome.cn:/home/simon/myquant/djapi/
|
||||
|
||||
# Single file sync MUST use full target path
|
||||
rsync -avz api/views.py simon@doorcome.cn:/home/simon/myquant/djapi/api/views.py
|
||||
```
|
||||
|
||||
## API Docs
|
||||
- /api/docs/ — Swagger UI
|
||||
- /api/redoc/ — ReDoc
|
||||
- /api/schema/ — OpenAPI JSON
|
||||
|
||||
## Video Processing
|
||||
```bash
|
||||
cd api/video
|
||||
python main.py
|
||||
```
|
||||
@@ -1,120 +0,0 @@
|
||||
# the name by which the project can be referenced within Serena
|
||||
project_name: "djapi"
|
||||
|
||||
|
||||
# list of languages for which language servers are started; choose from:
|
||||
# al ansible bash clojure cpp
|
||||
# cpp_ccls crystal csharp csharp_omnisharp dart
|
||||
# elixir elm erlang fortran fsharp
|
||||
# go groovy haskell haxe hlsl
|
||||
# java json julia kotlin lean4
|
||||
# lua luau markdown matlab msl
|
||||
# nix ocaml pascal perl php
|
||||
# php_phpactor powershell python python_jedi python_ty
|
||||
# r rego ruby ruby_solargraph rust
|
||||
# scala solidity swift systemverilog terraform
|
||||
# toml typescript typescript_vts vue yaml
|
||||
# zig
|
||||
# (This list may be outdated. For the current list, see values of Language enum here:
|
||||
# https://github.com/oraios/serena/blob/main/src/solidlsp/ls_config.py
|
||||
# For some languages, there are alternative language servers, e.g. csharp_omnisharp, ruby_solargraph.)
|
||||
# Note:
|
||||
# - For C, use cpp
|
||||
# - For JavaScript, use typescript
|
||||
# - For Free Pascal/Lazarus, use pascal
|
||||
# Special requirements:
|
||||
# Some languages require additional setup/installations.
|
||||
# See here for details: https://oraios.github.io/serena/01-about/020_programming-languages.html#language-servers
|
||||
# When using multiple languages, the first language server that supports a given file will be used for that file.
|
||||
# The first language is the default language and the respective language server will be used as a fallback.
|
||||
# Note that when using the JetBrains backend, language servers are not used and this list is correspondingly ignored.
|
||||
languages:
|
||||
- typescript
|
||||
- python
|
||||
|
||||
# the encoding used by text files in the project
|
||||
# For a list of possible encodings, see https://docs.python.org/3.11/library/codecs.html#standard-encodings
|
||||
encoding: "utf-8"
|
||||
|
||||
# line ending convention to use when writing source files.
|
||||
# Possible values: unset (use global setting), "lf", "crlf", or "native" (platform default)
|
||||
# This does not affect Serena's own files (e.g. memories and configuration files), which always use native line endings.
|
||||
line_ending:
|
||||
|
||||
# The language backend to use for this project.
|
||||
# If not set, the global setting from serena_config.yml is used.
|
||||
# Valid values: LSP, JetBrains
|
||||
# Note: the backend is fixed at startup. If a project with a different backend
|
||||
# is activated post-init, an error will be returned.
|
||||
language_backend:
|
||||
|
||||
# whether to use project's .gitignore files to ignore files
|
||||
ignore_all_files_in_gitignore: true
|
||||
|
||||
# advanced configuration option allowing to configure language server-specific options.
|
||||
# Maps the language key to the options.
|
||||
# Have a look at the docstring of the constructors of the LS implementations within solidlsp (e.g., for C# or PHP) to see which options are available.
|
||||
# No documentation on options means no options are available.
|
||||
ls_specific_settings: {}
|
||||
|
||||
# list of additional paths to ignore in this project.
|
||||
# Same syntax as gitignore, so you can use * and **.
|
||||
# Note: global ignored_paths from serena_config.yml are also applied additively.
|
||||
ignored_paths: []
|
||||
|
||||
# whether the project is in read-only mode
|
||||
# If set to true, all editing tools will be disabled and attempts to use them will result in an error
|
||||
# Added on 2025-04-18
|
||||
read_only: false
|
||||
|
||||
# list of tool names to exclude.
|
||||
# This extends the existing exclusions (e.g. from the global configuration)
|
||||
# Find the list of tools here: https://oraios.github.io/serena/01-about/035_tools.html
|
||||
excluded_tools: []
|
||||
|
||||
# list of tools to include that would otherwise be disabled (particularly optional tools that are disabled by default).
|
||||
# This extends the existing inclusions (e.g. from the global configuration).
|
||||
# Find the list of tools here: https://oraios.github.io/serena/01-about/035_tools.html
|
||||
included_optional_tools: []
|
||||
|
||||
# fixed set of tools to use as the base tool set (if non-empty), replacing Serena's default set of tools.
|
||||
# This cannot be combined with non-empty excluded_tools or included_optional_tools.
|
||||
# Find the list of tools here: https://oraios.github.io/serena/01-about/035_tools.html
|
||||
fixed_tools: []
|
||||
|
||||
# list of mode names that are to be activated by default, overriding the setting in the global configuration.
|
||||
# The full set of modes to be activated is base_modes (from global config) + default_modes + added_modes.
|
||||
# If the setting is undefined/empty, the default_modes from the global configuration (serena_config.yml) apply.
|
||||
# Otherwise, this overrides the setting from the global configuration (serena_config.yml).
|
||||
# Therefore, you can set this to [] if you do not want the default modes defined in the global config to apply
|
||||
# for this project.
|
||||
# This setting can, in turn, be overridden by CLI parameters (--mode).
|
||||
# See https://oraios.github.io/serena/02-usage/050_configuration.html#modes
|
||||
default_modes:
|
||||
|
||||
# list of mode names to be activated additionally for this project, e.g. ["query-projects"]
|
||||
# The full set of modes to be activated is base_modes (from global config) + default_modes + added_modes.
|
||||
# See https://oraios.github.io/serena/02-usage/050_configuration.html#modes
|
||||
added_modes:
|
||||
|
||||
# initial prompt for the project. It will always be given to the LLM upon activating the project
|
||||
# (contrary to the memories, which are loaded on demand).
|
||||
initial_prompt: ""
|
||||
|
||||
# time budget (seconds) per tool call for the retrieval of additional symbol information
|
||||
# such as docstrings or parameter information.
|
||||
# This overrides the corresponding setting in the global configuration; see the documentation there.
|
||||
# If null or missing, use the setting from the global configuration.
|
||||
symbol_info_budget:
|
||||
|
||||
# list of regex patterns which, when matched, mark a memory entry as read‑only.
|
||||
# Extends the list from the global configuration, merging the two lists.
|
||||
read_only_memory_patterns: []
|
||||
|
||||
# list of regex patterns for memories to completely ignore.
|
||||
# Matching memories will not appear in list_memories or activate_project output
|
||||
# and cannot be accessed via read_memory or write_memory.
|
||||
# To access ignored memory files, use the read_file tool on the raw file path.
|
||||
# Extends the list from the global configuration, merging the two lists.
|
||||
# Example: ["_archive/.*", "_episodes/.*"]
|
||||
ignored_memory_patterns: []
|
||||
+1
-1
@@ -99,7 +99,7 @@ python api/video/main.py
|
||||
|
||||
## 注意事项
|
||||
|
||||
- `config.py` 中的 TS_TOKEN 和 `deepseek.py`/`ai.py`/`audioRead.py` 中的 API key、`mysqlHandle.py` 中的数据库密码均为硬编码 —— 生产环境应迁移到环境变量
|
||||
- 所有密钥已迁移到环境变量,通过 `.env` 统一管理
|
||||
- `api/stock/` 下的模块支持两种导入方式(相对导入和绝对导入),这是为了兼容「作为 Django app 被调用」和「直接命令行运行脚本」两种场景
|
||||
- `api/video/` 模块设计为独立命令行运行,不依赖 Django 框架
|
||||
- `db.sqlite3` 已提交到代码库,包含 Django admin 的用户数据
|
||||
|
||||
+1
-1
@@ -98,7 +98,7 @@ uwsgi --stop uwsgi.pid
|
||||
| `news/reports/` | report_type, start_date, end_date, id | 日报查询(默认最近 24h;传 id 返回单份详情含事件) |
|
||||
| `news/events/` | days, importance, report_type, section, limit | 重要事件聚合(跨日报,最近 N 天 importance≥阈值) |
|
||||
|
||||
日报查询接口详细说明见 [`docs/news_report_api.md`](../docs/news_report_api.md)(表结构见 `djapi/docs/db_schema.md`)。
|
||||
日报查询接口详细说明见 [`docs/news_report_api.md`](../docs/news_report_api.md)(表结构见 `../docs/db_schema_v1.1.md`)。
|
||||
|
||||
API 文档(Swagger):`/api/docs/`
|
||||
OpenAPI Schema:`/api/schema/`
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
"""
|
||||
news_report / news_event 只读查询层(日报结构化入库,见 docs/db_schema.md)。
|
||||
news_report / news_event 只读查询层(日报结构化入库,见 docs/db_schema_v1.1.md)。
|
||||
|
||||
连接配置来自环境变量(与 docs/report_db_design.md §7 保持一致):
|
||||
连接配置来自环境变量(与 docs/db_schema_v1.1.md 保持一致):
|
||||
NEWS_DB_HOST / NEWS_DB_PORT / NEWS_DB_USER / NEWS_DB_PASSWORD / NEWS_DB_NAME
|
||||
NEWS_DB_PASSWORD 缺失时直接报错,禁止默认密码。
|
||||
|
||||
@@ -45,6 +45,18 @@ def _connect():
|
||||
return mysql.connector.connect(**load_db_config())
|
||||
|
||||
|
||||
def _parse_json(value):
|
||||
"""把 JSON 字符串列(如 sources)解析为 dict/list;已是对象则原样返回。"""
|
||||
if value is None:
|
||||
return None
|
||||
if isinstance(value, str):
|
||||
try:
|
||||
return json.loads(value)
|
||||
except (TypeError, ValueError):
|
||||
return None
|
||||
return value
|
||||
|
||||
|
||||
def _row_to_dict(row: dict) -> dict:
|
||||
"""序列化行:stats JSON 解析、日期/时间转 ISO 字符串。"""
|
||||
d = dict(row)
|
||||
@@ -87,11 +99,14 @@ def fetch_reports(
|
||||
report = _row_to_dict(row)
|
||||
cur.execute(
|
||||
"SELECT id, section, rank, importance, event_type, title, "
|
||||
"summary, sentiment, source, url "
|
||||
"summary, sentiment, source, sources, url "
|
||||
"FROM news_event WHERE report_id = %s ORDER BY section, rank",
|
||||
(report_id,),
|
||||
)
|
||||
report["events"] = [dict(r) for r in cur.fetchall()]
|
||||
report["events"] = [
|
||||
{**dict(r), "sources": _parse_json(r["sources"])}
|
||||
for r in cur.fetchall()
|
||||
]
|
||||
return report
|
||||
|
||||
where, params = [], []
|
||||
@@ -153,7 +168,7 @@ def fetch_important_events(
|
||||
sql = (
|
||||
"SELECT r.report_date, r.report_type, e.id, e.section, e.rank, "
|
||||
"e.importance, e.event_type, e.title, e.summary, e.sentiment, "
|
||||
"e.source, e.url "
|
||||
"e.source, e.sources, e.url "
|
||||
"FROM news_event e "
|
||||
"JOIN news_report r ON r.id = e.report_id "
|
||||
"WHERE " + " AND ".join(where)
|
||||
@@ -163,6 +178,10 @@ def fetch_important_events(
|
||||
)
|
||||
params.append(int(limit))
|
||||
cur.execute(sql, tuple(params))
|
||||
return [dict(r) for r in cur.fetchall()]
|
||||
rows = cur.fetchall()
|
||||
return [
|
||||
{**dict(r), "sources": _parse_json(r["sources"])}
|
||||
for r in rows
|
||||
]
|
||||
finally:
|
||||
conn.close()
|
||||
|
||||
@@ -14,6 +14,7 @@ class EventSerializer(serializers.Serializer):
|
||||
summary = serializers.CharField(allow_null=True)
|
||||
sentiment = serializers.CharField(allow_null=True)
|
||||
source = serializers.CharField(allow_null=True)
|
||||
sources = serializers.JSONField(allow_null=True)
|
||||
url = serializers.CharField(allow_null=True)
|
||||
|
||||
|
||||
@@ -47,4 +48,5 @@ class ImportantEventSerializer(serializers.Serializer):
|
||||
summary = serializers.CharField(allow_null=True)
|
||||
sentiment = serializers.CharField(allow_null=True)
|
||||
source = serializers.CharField(allow_null=True)
|
||||
sources = serializers.JSONField(allow_null=True)
|
||||
url = serializers.CharField(allow_null=True)
|
||||
|
||||
@@ -122,6 +122,17 @@ class NewsEventsAPITest(TestCase):
|
||||
self.assertEqual(kwargs['report_type'], 'intl')
|
||||
self.assertEqual(kwargs['section'], 'intl')
|
||||
|
||||
@patch('api.report.query.fetch_important_events',
|
||||
return_value=[{'id': 1, 'title': 'x', 'source': 'yicai',
|
||||
'sources': ['yicai', 'stcn']}])
|
||||
def test_sources_field_passthrough(self, mock_fetch):
|
||||
"""事件聚合响应原样透传 sources(JSON 数组)"""
|
||||
resp = self.client.get(self.url)
|
||||
self.assertEqual(resp.status_code, 200)
|
||||
body = resp.json()
|
||||
self.assertEqual(body[0]['sources'], ['yicai', 'stcn'])
|
||||
self.assertEqual(body[0]['source'], 'yicai')
|
||||
|
||||
def test_invalid_days(self):
|
||||
resp = self.client.get(self.url, {'days': 'abc'})
|
||||
self.assertEqual(resp.status_code, 400)
|
||||
|
||||
@@ -1,135 +0,0 @@
|
||||
# continuation.md
|
||||
|
||||
## 当前项目状态
|
||||
|
||||
djapi — Django 5.2 金融数据 API 项目,2026-06-17 已部署。
|
||||
|
||||
## Checkpoint 记录
|
||||
|
||||
| 日期 | 内容 |
|
||||
|------|------|
|
||||
| 2026-08-03 | 新增日报查询 API ×2(news/reports/ + news/events/),基于 news_report/news_event 表,已部署 doorcome ✅ |
|
||||
|
||||
服务器:`simon@doorcome.cn`,路径 `/home/simon/myquant/djapi/`,虚拟环境 `/opt/miniconda/envs/django/`。
|
||||
|
||||
## 已完成
|
||||
|
||||
### 1. 安全:密钥统一管理
|
||||
- 所有密钥 → 环境变量,`.env` 统一管理
|
||||
- `djapi/env_loader.py`(Django 端)+ `api/video/env.py`(video 端)双加载器
|
||||
- 共享 `MySQLDB` → `api/utils/mysql_handler.py`
|
||||
- `.env.example`、`.gitignore`
|
||||
|
||||
### 2. 代码质量
|
||||
- `api/views.py`:227 → ~130 行,消除重复
|
||||
- `api/stock/stock_utils.py`:`viewFunc_singleParam()` 包装器
|
||||
- `api/stock/config.py`:拆分为 config / strategy_config / scan_config
|
||||
|
||||
### 3. drf-spectacular 集成
|
||||
- 14 端点 `@api_view` + `@extend_schema`,8 tag 分组
|
||||
- 13 Serializer,Swagger `/api/docs/`
|
||||
|
||||
### 4. 股息率 API 优化
|
||||
- **删除** `api/stock/getDivData_AK.py`(akshare 版),`/api/getdivak/` 路由移除
|
||||
- **优化** `api/stock/getStockDiv2.py`:
|
||||
- TTM 计算:`calculate_ttm_div` 行级循环 O(n²) → `rolling('360D').sum()` O(n)
|
||||
- 删除向前填充逻辑(~30 行),避免与毛刺平滑冲突
|
||||
- **修复** `api/stock/smoothBrush.py`:if/elif 分支中 prev_valid/next_valid 赋值反了
|
||||
|
||||
### 5. video 模块重构与 Bug 修复
|
||||
- 新增 `api/video/env.py` — .env 加载
|
||||
- `newsRedo.py` 重写 — 三分支智能重处理
|
||||
- **P0 修复**:`getVideo5.py` 日期校验 bug(`start_date > start_date` → `start_date > end_date`)
|
||||
- **P1 清理**:删除 `ai.py`(两个函数均为死代码),清理 `newsProcess.py` 冗余 import
|
||||
- **P2 修复**:`audioRead.py` — `transcribe_audio` 异常时返回 `['', '']` 统一类型;`analyze_and_correct_text` 防御 None
|
||||
- **P3 修复**:`newsProcess.py` — `news_to_db()` JSON 解析自适应 dict/list(DeepSeek json_object 模式返回 dict 包装)
|
||||
|
||||
### 6. 文档与测试
|
||||
- `CLAUDE.md`、`README.md`、`continuation.md`
|
||||
- 18 个单元测试
|
||||
|
||||
## video 目录文件现状(10 个 .py)
|
||||
|
||||
| 文件 | 职责 |
|
||||
|------|------|
|
||||
| `env.py` | .env 加载 |
|
||||
| `getVideo5.py` | 主流程:抓取→下载→ASR→入库 |
|
||||
| `audioRead.py` | 音频转换、分割、ASR 识别、文本纠错 |
|
||||
| `deepseek.py` | DeepSeek API 封装(类 + `deepseek_text` 函数,支持 `response_format`) |
|
||||
| `newsProcess.py` | AI 新闻分割+标题提取,JSON 自适应解析 |
|
||||
| `newsRedo.py` | 手动重处理(三分支) |
|
||||
| `main.py` | 定时任务入口(当天) |
|
||||
| `main_videos.py` | 批量补缺(扫描缺失日期) |
|
||||
| `mysqlHandle.py` | MySQLDB 重新导出 |
|
||||
|
||||
## 所有 API 端点(16 个)
|
||||
|
||||
| 端点 | 数据源 | 说明 |
|
||||
|------|--------|------|
|
||||
| `stockbasic/` | Tushare | 日线行情 |
|
||||
| `stockinfo/` | Tushare | 个股基本信息 |
|
||||
| `industrys/` | Tushare | 行业股票列表 |
|
||||
| `stockparam/` | Tushare | 个股参数 |
|
||||
| `stockep/` | Tushare | TTM EPS |
|
||||
| `quarterlyEps/` | Tushare | 季度 EPS |
|
||||
| `indexByName/` | Tushare | 指数查询 |
|
||||
| `indexDatas/` | Tushare | 指数行情 |
|
||||
| `dailymargin/` | Tushare | 每日融资融券汇总 |
|
||||
| `stockmargin/` | Tushare | 个股融资融券 |
|
||||
| `finance/` | Tushare | 财务报表分析 |
|
||||
| `getdiv/` | Tushare | 股息率(TTM rolling + 毛刺平滑) |
|
||||
| `xwlbNews/` | MySQL | 新闻联播原始文本 |
|
||||
| `xwlbFine/` | MySQL | 新闻联播 AI 精编 |
|
||||
| `news/reports/` | MySQL (news_) | 日报查询:默认最近 24h;传 id 返回详情含事件 |
|
||||
| `news/events/` | MySQL (news_) | 重要事件聚合:最近 N 天 importance≥阈值 |
|
||||
|
||||
---
|
||||
|
||||
## 日报查询 API(2026-08-03 新增)
|
||||
|
||||
### 模块
|
||||
|
||||
- 新增 `api/report/` 包(独立于 stock):`query.py`(连库+查询 SQL)/ `views.py`(2 视图)/ `serializers.py`(OpenAPI)/ `tests.py`(17 个单测,mock 查询层)
|
||||
- `api/urls.py` 注册 `news/reports/`、`news/events/`;`settings.py` SPECTACULAR TAGS 加「日报」
|
||||
- 数据库:doorcome 本机 MariaDB `myquant` 库 `news_report`(180 行)+ `news_event`(4372 条),与现有 `MYSQL_*` 同库同用户
|
||||
- 连接配置:服务器 `djapi/.env` 新增 `NEWS_DB_*`(复用 MYSQL_* 值,密码必填否则 500)
|
||||
- 文档:`docs/news_report_api.md`(使用手册,含线上地址/curl/真实样例);README API 概览表已加两行
|
||||
|
||||
### 部署(2026-08-03 完成)
|
||||
|
||||
- rsync 增量同步(**未用文档中的 --delete**,见下)→ 重启 uWSGI → 冒烟通过(列表/详情/聚合/400/404)
|
||||
- 线上:`https://api.doorcome.cn/api/news/reports/`、`/api/news/events/`,Swagger `/api/docs/`「日报」tag
|
||||
|
||||
### 已知事项
|
||||
|
||||
1. **服务器顶层历史平铺文件未清理**:`/home/simon/myquant/djapi/` 顶层有 views.py/urls.py/smoothBrush.py/getStockDiv2.py/env_loader.py/akshare_data.py(历史 rsync 陷阱产物),`--delete` 会删除它们,但 `divSearch.py`(离线脚本)仍绝对导入顶层 getStockDiv2/smoothBrush → 本次增量同步保留;清理前需先修 divSearch.py 的导入
|
||||
2. **既有失败测试**:`api.tests.DateFormatCorrectionTest.test_empty_string`(date_format_correction('') 期望 None 实得 ''),与本次无关
|
||||
3. 冒烟曾发现 fetch_reports 列表 SQL 缺 `r.` 别名前缀(1052 ambiguous),已修复
|
||||
4. 本地验证需绕过 macOS TCC:`HOME=/tmp/djtest_home PYTHONPATH=/tmp/djtest_pkgs`(tushare 写 ~/tk.csv 被拦 + quant 环境缺 mysql-connector-python)
|
||||
|
||||
## 部署
|
||||
|
||||
```bash
|
||||
# 全量同步
|
||||
rsync -avz --delete \
|
||||
--exclude='.env' --exclude='db.sqlite3' \
|
||||
--exclude='*.log' --exclude='uwsgi.pid' \
|
||||
--exclude='__pycache__/' --exclude='*.pyc' \
|
||||
--exclude='xwlb_video/' --exclude='audio_processing/' \
|
||||
/Users/summer/Downloads/cc-cursor/djapi/ \
|
||||
simon@doorcome.cn:/home/simon/myquant/djapi/
|
||||
|
||||
# 单文件同步必须写完整路径
|
||||
# 正确:rsync api/views.py simon@...:/.../djapi/api/views.py
|
||||
|
||||
# 重启
|
||||
ssh simon@doorcome.cn "kill \$(lsof -ti:5004); sleep 2; /opt/miniconda/envs/django/bin/uwsgi --ini /home/simon/myquant/djapi/uwsgi.ini"
|
||||
```
|
||||
|
||||
## 关键设计决策
|
||||
|
||||
- video 模块保护、向后兼容优先、不使用 python-dotenv
|
||||
- rsync 陷阱:多文件源会展平路径
|
||||
- `.env` 双加载:Django 端 `djapi/env_loader.py` + video 端 `api/video/env.py`
|
||||
- 股息率 TTM 用 `rolling('360D').sum()` 向量化,不手动循环
|
||||
- DeepSeek json_object 模式返回 dict,newsProcess 自适应提取 list
|
||||
@@ -1,330 +0,0 @@
|
||||
# Milestone 10 后端实现逻辑:日报结构化入库
|
||||
|
||||
> 版本:v0.1(设计稿) | 2026-07
|
||||
> 对应 project_plan.md「十八、Milestone 10」
|
||||
> **范围**:本项目侧"后端"= 数据生产层(日报内容生成 + 结构化写入 MySQL)。
|
||||
> 不包含 API 服务与前端页面(由用户另行实现),但表结构与数据契约以本文档为准,供 API/前端对接。
|
||||
|
||||
---
|
||||
|
||||
## 1. 定位
|
||||
|
||||
现有链路:`reporter.py` 收集数据 → `_render_html()` 渲染 HTML → scp 上传 doorcome。
|
||||
改造后:`reporter.py` 收集数据 → 组装结构化 `ReportData` → 写入 MySQL(`news_report` / `news_event`),不再产出 HTML。
|
||||
|
||||
另需:把 doorcome 上 178 份历史日报 HTML(`finance_news_daily_*` ×50、`intl_news_daily_*` ×128)解析成同一 `ReportData` 结构入库。
|
||||
|
||||
---
|
||||
|
||||
## 2. 数据流总览
|
||||
|
||||
```
|
||||
[历史 HTML ×178] [每日 pipeline]
|
||||
doorcome:/var/www/html/echart/research/ crawler→extractor→dedup→llm→embed→qdrant
|
||||
(一次性 scp 到 data/reports_history/) │
|
||||
│ ▼
|
||||
▼ reporter.generate_report()
|
||||
report_import/parser.py │
|
||||
│ (BeautifulSoup 解析) ▼
|
||||
▼ 组装 ReportData 组装 ReportData
|
||||
report_import/importer.py │
|
||||
│ (幂等 upsert) ▼
|
||||
▼ │
|
||||
┌────────────────────── MySQL (myquant 库) ──────────────────────┐
|
||||
│ news_report(主表) news_event(事件明细) │
|
||||
└────────────────────────────────────────────────────────────────┘
|
||||
▲
|
||||
API / 前端(用户另行实现,只读)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. 数据模型(Pydantic,`report_db/models.py`)
|
||||
|
||||
```python
|
||||
class EventRow(BaseModel):
|
||||
"""一条事件记录,对应 news_event 一行。"""
|
||||
section: str # xwlb | news | cninfo | intl
|
||||
rank: int # 板块内序号(从 1 开始)
|
||||
importance: int | None = None
|
||||
event_type: str | None = None
|
||||
title: str
|
||||
summary: str | None = None
|
||||
sentiment: str | None = None # positive | negative | neutral | ''
|
||||
source: str | None = None # 来源(如 cls / ForexLive)
|
||||
url: str | None = None
|
||||
|
||||
class ReportData(BaseModel):
|
||||
"""一份完整日报,对应 news_report 一行 + news_event 多行。"""
|
||||
report_date: date # 日报日期(YYYY-MM-DD)
|
||||
report_type: str # finance | intl
|
||||
file_name: str # 源文件名(新生成时可为 "")
|
||||
generated_at: datetime # 生成时间
|
||||
ai_summary: str | None = None
|
||||
stats: dict[str, Any] = Field(default_factory=dict) # 数据总览统计快照 → JSON 列
|
||||
events: list[EventRow] = Field(default_factory=list)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. 字段映射(核心契约)
|
||||
|
||||
### 4.1 事件 JSON(data/events/)→ news_event
|
||||
|
||||
现有事件文件结构与 news_event 字段对应关系(reporter 收集时直接转换):
|
||||
|
||||
| news_event 字段 | 事件 JSON 来源 |
|
||||
| --- | --- |
|
||||
| section | 来源判定:`source_id=="cninfo"` → `cninfo`;`source_id=="xwlb"` → `xwlb`;否则 `news`;intl 解析固定 `intl` |
|
||||
| importance | `event.importance` |
|
||||
| event_type | `event.event_type` |
|
||||
| title | `title` |
|
||||
| summary | `event.summary` |
|
||||
| sentiment | `event.sentiment` |
|
||||
| source | `source_id` |
|
||||
| url | `url`(xwlb 为空) |
|
||||
|
||||
### 4.2 历史 HTML → ReportData
|
||||
|
||||
解析策略:**表头驱动列映射**。不同日报表格列集合不同:
|
||||
|
||||
| 板块 | 表格列(<th>) | section |
|
||||
| --- | --- | --- |
|
||||
| 新闻联播(finance) | `# / (空) / 标题 / 重要度 / 事件类型` | xwlb |
|
||||
| 重要事件:新闻(finance) | `# / (空) / 标题 / 源 / 重要度 / 事件类型 / 摘要` | news |
|
||||
| 重要事件:公告调研(finance) | 同上 | cninfo |
|
||||
| 重要事件(intl) | `# / (空) / 标题 / 重要度 / 事件类型 / 摘要` | intl |
|
||||
|
||||
要点:
|
||||
|
||||
- 以表头文本定位列索引("标题""重要度""事件类型""摘要""源"),空 `<th>` 为情绪图标列(⚪/🔴/🟢 → neutral/negative/positive),**不要依赖列位置**。
|
||||
- 情绪图标仅存在于有图标列的表;intl 表情绪列存在,finance 表情绪列存在(空 th 首列后)。
|
||||
- intl 无"源"列时,尝试从标题尾部 `[来源]` 或摘要尾部提取,提取不到则 `source=None`。
|
||||
- 标题中的股票代码标注 `(600519, ...)` 与 ⭐(自选股标记)需剥除,只保留纯标题。
|
||||
- AI 摘要:取 `h2`("一、AI 摘要")之后紧随的 `.ai-summary` 区块纯文本(保留换行)。
|
||||
- 数据总览 → `stats` JSON:按 `h3` 标题映射 key(见 4.3),解析该 h3 后的首个 `<table>`,缺失的板块跳过、不报错。
|
||||
- 容错:任一板块解析失败 → 记 WARNING 日志,该板块置空,不影响整份入库;整份文件解析失败 → 抛 `ReportParseError`(由 importer 捕获计数)。
|
||||
|
||||
### 4.3 数据总览 → stats JSON
|
||||
|
||||
| h3 标题(含板块名) | stats key |
|
||||
| --- | --- |
|
||||
| M1→M6 管道 / 管道 | `pipeline`(保留原始行) |
|
||||
| 各源数据 | `sources` |
|
||||
| 情绪分布 | `sentiment` |
|
||||
| 重要度分布 | `importance` |
|
||||
| 事件类型(TOP 10 / 分布) | `event_types` |
|
||||
| 文章来源分布 | `source_dist` |
|
||||
|
||||
`stats` 存 MySQL `JSON` 列,前端自行解析展示。历史文件与未来新日报的 stats 结构可能不同(finance 与 intl 板块不同),一律按快照存储,不做跨版本规范化。
|
||||
|
||||
---
|
||||
|
||||
## 5. 模块设计
|
||||
|
||||
### 5.1 新包 `report_db/`(DB 层)
|
||||
|
||||
```
|
||||
report_db/
|
||||
├── __init__.py # 导出 connect / init_schema / save_report
|
||||
├── models.py # EventRow / ReportData(Pydantic)
|
||||
├── schema.py # DDL 常量(news_report / news_event,见 project_plan.md 十八)
|
||||
└── db.py # 连接、事务、写入
|
||||
```
|
||||
|
||||
`db.py` 关键函数:
|
||||
|
||||
```python
|
||||
def load_db_config() -> DbConfig:
|
||||
"""从环境变量读取 NEWS_DB_HOST/PORT/USER/PASSWORD/NAME。
|
||||
缺失 PASSWORD 时记 ERROR 并 raise,禁止默认密码。"""
|
||||
|
||||
def connect(cfg: DbConfig) -> Connection:
|
||||
"""pymysql.connect(autocommit=False, charset="utf8mb4", cursorclass=DictCursor)。
|
||||
失败时 logger.exception + raise。"""
|
||||
|
||||
def init_schema(conn: Connection) -> None:
|
||||
"""执行 schema.py 中的 CREATE TABLE IF NOT EXISTS ×2。"""
|
||||
|
||||
def save_report(conn: Connection, report: ReportData) -> int:
|
||||
"""事务内:
|
||||
1. INSERT INTO news_report (...) VALUES (...) 或按 (report_date, report_type, file_name)
|
||||
唯一键命中时 UPDATE(新生成日报重复执行 = 覆盖同 file_name/同日期,幂等);
|
||||
2. 取 report_id,DELETE 旧事件后批量 INSERT news_event(保证整份覆盖一致)。
|
||||
返回 report_id。"""
|
||||
|
||||
def transaction(conn: Connection) -> contextmanager:
|
||||
"""提交/回滚上下文管理器。"""
|
||||
|
||||
def fetch_report(conn: Connection, report_id: int) -> dict | None:
|
||||
"""读侧辅助(联调/测试用),API 侧由用户自行实现。"""
|
||||
```
|
||||
|
||||
要点:
|
||||
|
||||
- 所有 SQL 为 MySQL/MariaDB 方言(`JSON` 列、`ENGINE=InnoDB`、`COMMENT`),**不依赖 ORM**。
|
||||
- 连接生命周期:每次 `save_report` 短连接(report 一天跑几次,量小,无需连接池;如未来加大再换)。
|
||||
- 字符集 utf8mb4,`SET NAMES utf8mb4` 由 pymysql charset 参数处理。
|
||||
|
||||
### 5.2 新包 `report_import/`(历史解析)
|
||||
|
||||
```
|
||||
report_import/
|
||||
├── __init__.py
|
||||
├── parser.py # parse_finance_report / parse_intl_report(BeautifulSoup)
|
||||
└── importer.py # import_history(dir, date=None, type=None) -> ImportStats
|
||||
```
|
||||
|
||||
`parser.py`:
|
||||
|
||||
```python
|
||||
class ReportParseError(Exception): ...
|
||||
|
||||
def parse_finance_report(html: str, file_name: str) -> ReportData: ...
|
||||
def parse_intl_report(html: str, file_name: str) -> ReportData: ...
|
||||
def parse_report(html: str, file_name: str) -> ReportData:
|
||||
"""按文件名前缀分流:finance_news_daily_* / intl_news_daily_*。"""
|
||||
```
|
||||
|
||||
- 依赖复用现有 `beautifulsoup4`(已在 pyproject 依赖),**不新增解析库**。
|
||||
- `report_date` 从文件名解析(`*_daily_{YYYYMMDD}_*.html`),不信任目录名。
|
||||
- `generated_at` 从文件名时间(`{HHMMSS}`)或 `<header>` 中"生成于"文本解析,解析不到用文件 mtime。
|
||||
|
||||
`importer.py`:
|
||||
|
||||
```python
|
||||
@dataclass
|
||||
class ImportStats:
|
||||
scanned: int = 0 # 扫描到的日报文件数
|
||||
imported: int = 0 # 新入库
|
||||
skipped: int = 0 # 已存在(幂等跳过)
|
||||
failed: int = 0 # 解析失败
|
||||
errors: list[str] = field(default_factory=list)
|
||||
|
||||
def import_history(report_dir: Path, date: str | None = None,
|
||||
report_type: str | None = None) -> ImportStats:
|
||||
"""遍历 {report_dir}/{YYYYMMDD}/*_news_daily_*.html,
|
||||
过滤 date / type,逐个 parse → save_report。"""
|
||||
```
|
||||
|
||||
### 5.3 `scheduler/reporter.py` 改造(完全切换)
|
||||
|
||||
- 新增 `_build_report_data(news, cninfo, pipeline, ai_summary, day_str, xwlb) -> ReportData`:
|
||||
- 事件转换:`news["high"]` → `EventRow(section="news", ...)`;`cninfo["high"]` → `section="cninfo"`;`xwlb["items"]` → `section="xwlb"`;
|
||||
- `stats` 组装:`{"pipeline": pipeline, "sources": {...}, "sentiment": news["sentiments"], "importance": news["importances"], "event_types": news["event_types"], "cninfo": {...}}`;
|
||||
- 事件 `rank` 按板块内顺序编号。
|
||||
- `generate_report(day_str, *, upload=True)` 改为:收集(逻辑不变)→ `_build_report_data` → `connect()` + `save_report()`;删除 `_render_html`/`_upload` 调用。
|
||||
- `_render_*` 函数**保留但标记 deprecated**(注释说明"完全切换后不再调用"),不删除,保证最小改动、可回退。
|
||||
- 返回值由 `Path | None` 改为 `report_id: int | None`;`scheduler/pipeline.py` 中 report 步骤仅判断非 None(实施时核实该处调用,保持兼容)。
|
||||
- `stock_reporter.py` **不改动**(个股日报不在本期范围)。
|
||||
|
||||
### 5.4 `a_share_cli/main.py` 新增子命令
|
||||
|
||||
```
|
||||
uv run a-share report-import [--dir data/reports_history] [--date YYYYMMDD] [--type finance|intl]
|
||||
```
|
||||
|
||||
- 默认全量扫描 `REPORT_HISTORY_DIR`(.env 可配,默认 `data/reports_history/`)。
|
||||
- 输出 ImportStats 汇总(扫描/导入/跳过/失败)。
|
||||
|
||||
---
|
||||
|
||||
## 6. 关键流程
|
||||
|
||||
### 6.1 历史导入(一次执行,可重复)
|
||||
|
||||
```
|
||||
1. scp -r doorcome:/var/www/html/echart/research/2026* → data/reports_history/
|
||||
(一次手工操作,不进代码)
|
||||
2. uv run a-share report-import
|
||||
for each {date}/{file}:
|
||||
report_type = 文件名前缀(finance|intl)
|
||||
ReportData = parse_report(html, file_name)
|
||||
try: save_report(conn, ReportData) → imported += 1
|
||||
except DuplicateKey: skipped += 1 # 已导入过
|
||||
except ReportParseError as e: failed += 1; errors.append(str(e))
|
||||
3. 校验: SELECT report_type, COUNT(*) FROM news_report GROUP BY report_type
|
||||
期望 50 / 128
|
||||
```
|
||||
|
||||
### 6.2 每日日报生成(pipeline 07:00 步骤)
|
||||
|
||||
```
|
||||
generate_report(day_str):
|
||||
news = _collect_news_events(day_str) # 不变
|
||||
cninfo = _collect_cninfo_events(day_str) # 不变
|
||||
xwlb = _collect_xwlb(day_str) # 不变
|
||||
pipeline = _collect_pipeline_stats(day_str) # 不变
|
||||
ai_summary = _generate_ai_summary(...) # 不变
|
||||
report = _build_report_data(...) # 新增
|
||||
save_report(connect(), report) # 新增(替代渲染+上传)
|
||||
```
|
||||
|
||||
### 6.3 幂等策略
|
||||
|
||||
- 唯一键 `(report_date, report_type, file_name)`:
|
||||
- 历史导入:命中 → 跳过(或 `--force` 覆盖);
|
||||
- 新日报:`file_name=""` 时唯一键退化为 `(report_date, report_type, "")`,同一天重复跑 → UPDATE 覆盖,事件表 DELETE+INSERT 全量替换,**不产生历史残留**。
|
||||
|
||||
---
|
||||
|
||||
## 7. 配置项(.env / .env.example)
|
||||
|
||||
```env
|
||||
# ---- 日报结构化入库 (M10) ----
|
||||
NEWS_DB_HOST=127.0.0.1 # 开发走 ssh 隧道: ssh -L 13306:127.0.0.1:13306 pi
|
||||
NEWS_DB_PORT=13306
|
||||
NEWS_DB_USER=myquant
|
||||
NEWS_DB_PASSWORD= # 填真实值,禁止写入源码/文档
|
||||
NEWS_DB_NAME=myquant
|
||||
REPORT_HISTORY_DIR=data/reports_history
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 8. 错误处理
|
||||
|
||||
| 场景 | 行为 |
|
||||
| --- | --- |
|
||||
| DB 不可达/凭据错误 | `connect()` 抛异常 → `generate_report` 记 ERROR 并返回 None(pipeline 该步骤失败,其余步骤不受影响) |
|
||||
| 单份历史文件解析失败 | 记 WARNING,`failed += 1`,继续下一份;结束输出失败清单 |
|
||||
| 事件字段缺失(如无摘要列) | 对应字段留 None,不抛错 |
|
||||
| 全部失败 | `report-import` 返回非 0 退出码,便于排查 |
|
||||
|
||||
---
|
||||
|
||||
## 9. 测试策略(tests/)
|
||||
|
||||
| 文件 | 内容 |
|
||||
| --- | --- |
|
||||
| `tests/test_report_parser.py` | 用 fixtures(从 178 份中拷贝 finance/intl 各 1 份真实样例到 `tests/fixtures/`)断言:板块数、事件行数、字段映射、标题净化、幂等文件日期解析 |
|
||||
| `tests/test_report_db.py` | 纯逻辑:`_build_report_data` 组装正确;SQL 层用 sqlite3 内存库建同构(简化 DDL)验证 upsert/覆盖语义 |
|
||||
| `tests/test_report_import.py` | 临时目录构造 2-3 份假 HTML → 全流程导入 → 断言 ImportStats 计数与幂等 |
|
||||
| 集成(`@pytest.mark.integration`,默认跳过) | 连真实 MySQL:init_schema + save_report + 查询回读 |
|
||||
|
||||
新增 pytest marker 说明:真实 DB 连接一律走 integration,**单元测试不得依赖生产库**。
|
||||
|
||||
---
|
||||
|
||||
## 10. 依赖变更
|
||||
|
||||
- `uv add pymysql`(纯 Python 驱动,唯一新增依赖)
|
||||
- 解析复用现有 `beautifulsoup4`,不新增
|
||||
|
||||
---
|
||||
|
||||
## 11. 开放问题(沿自 project_plan.md 十八,不阻塞开发)
|
||||
|
||||
1. 生产连接:pi5 无法直连 `192.168.1.10:13306`(隧道仅绑 loopback)——需决定改 pi 的 autossh 绑定 / pi5 自建隧道。
|
||||
2. intl 日报生成方不在本项目,未来 intl 新日报需按同一表结构写入(本项目仅负责解析历史 + finance 新日报)。
|
||||
3. 个股日报(research 根目录文件)本期不处理。
|
||||
|
||||
---
|
||||
|
||||
## 12. 实施顺序(供开发排期)
|
||||
|
||||
1. `report_db/`(models/schema/db)+ `.env` 配置 + 建表验证
|
||||
2. `report_import/parser.py` + fixtures + 单测
|
||||
3. `report_import/importer.py` + CLI `report-import` + 178 份全量导入验收
|
||||
4. `reporter.py` 改造(_build_report_data + save_report)+ pipeline 兼容性验证
|
||||
5. docs/db_schema.md 定稿(给 API/前端)、README / continuation.md 更新
|
||||
@@ -1,4 +1,4 @@
|
||||
# CLAUDE.md
|
||||
# Agent 工作指南
|
||||
|
||||
cc-cursor — Mac Mini 单机量化研究平台。全链路:Data → Factor → Backtest → Optimize → ML → Sentiment → Agent。
|
||||
|
||||
@@ -6,11 +6,11 @@ cc-cursor — Mac Mini 单机量化研究平台。全链路:Data → Factor
|
||||
|
||||
| 模块 | 详情文件 | 核心入口 |
|
||||
|------|---------|---------|
|
||||
| 数据层 + 数据库 | `CLAUDE-data.md` | `from data.data_manager import DataManager` |
|
||||
| 因子引擎 + 情绪 | `CLAUDE-factors.md` | `from factors.registry import get_factor` |
|
||||
| 回测 + 优化 | `CLAUDE-backtest.md` | `from backtest.vectorbt.engine import VectorBTEngine` |
|
||||
| ML 模型 | `CLAUDE-ml.md` | `from models.lightgbm.model import LightGBMModel` |
|
||||
| Agent 系统 + CLI | `CLAUDE-agents.md` | `python finance/cli/agent_cli.py daily` |
|
||||
| 数据层 + 数据库 | [data-layer.md](data-layer.md) | `from data.data_manager import DataManager` |
|
||||
| 因子引擎 + 情绪 | [factors.md](factors.md) | `from factors.registry import get_factor` |
|
||||
| 回测 + 优化 | [backtest.md](backtest.md) | `from backtest.vectorbt.engine import VectorBTEngine` |
|
||||
| ML 模型 | [ml-models.md](ml-models.md) | `from models.lightgbm.model import LightGBMModel` |
|
||||
| Agent 系统 + CLI | [agents.md](agents.md) | `python finance/cli/agent_cli.py daily` | |
|
||||
|
||||
## 工作区布局
|
||||
|
||||
@@ -19,7 +19,7 @@ cc-cursor — Mac Mini 单机量化研究平台。全链路:Data → Factor
|
||||
| `finance/` | 核心量化引擎(**代码实际位置**)。代码内 import 用顶层名 `data.*`/`factors.*` 等 — 由 CLI 把 `finance/` 加入 sys.path;文件路径为 `finance/data/xxx.py` 等 |
|
||||
| `djapi/` | Django API 子项目,有独立 `djapi/CLAUDE.md` |
|
||||
| `shared/script/` | `autossh.sh` — MariaDB SSH 隧道 |
|
||||
| `docs/` | `usage.md` / `usage.html` 使用指南、`news_report_api.md` 新闻接口文档;**新建 md 一律放这里** |
|
||||
| `docs/` | 项目文档目录,包含使用指南、架构说明、API 参考等;**新建 md 一律放这里** |
|
||||
| `finance/strategy` `portfolio` `execution` `scheduler/` | 空壳占位(仅 `__init__.py`),逻辑未落地,别误以为有实现 |
|
||||
| `finance/reports/` | 日报输出 `daily_YYYYMMDD.{md,html}` |
|
||||
|
||||
@@ -36,13 +36,13 @@ bash shared/script/autossh.sh # DB SSH 隧道 (本地 13306 → 远程 3306)
|
||||
|
||||
| 任务类型 | 先读取 |
|
||||
|---------|--------|
|
||||
| 数据源/数据库/cache 相关 | `CLAUDE-data.md` |
|
||||
| 因子/情绪/新闻相关 | `CLAUDE-factors.md` + `CLAUDE-reference.md` |
|
||||
| 回测/优化/策略相关 | `CLAUDE-backtest.md` |
|
||||
| ML 模型/特征工程相关 | `CLAUDE-ml.md` |
|
||||
| Agent/CLI/报告相关 | `CLAUDE-agents.md` |
|
||||
| 第三方库 API/参数 | `web_fetch` / `research` 查官方文档 |
|
||||
| 因子名/类名/表结构速查 | `CLAUDE-reference.md` |
|
||||
| 数据源/数据库/cache 相关 | [data-layer.md](data-layer.md) |
|
||||
| 因子/情绪/新闻相关 | [factors.md](factors.md) + [reference.md](reference.md) |
|
||||
| 回测/优化/策略相关 | [backtest.md](backtest.md) |
|
||||
| ML 模型/特征工程相关 | [ml-models.md](ml-models.md) |
|
||||
| Agent/CLI/报告相关 | [agents.md](agents.md) |
|
||||
| 第三方库 API/参数 | 查官方文档 |
|
||||
| 因子名/类名/表结构速查 | [reference.md](reference.md) |
|
||||
|
||||
## 多步任务规则
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# CLAUDE-agents.md — Agent 系统 + CLI
|
||||
# Agent 系统 + CLI
|
||||
|
||||
## Agent 架构
|
||||
|
||||
+148
@@ -0,0 +1,148 @@
|
||||
# DJAPI — Django API 参考
|
||||
|
||||
Django 5.2 项目,提供 A 股金融数据 API 和新闻联播视频处理能力。部署在 Linux 服务器上,通过 uWSGI + nginx 对外服务。
|
||||
|
||||
---
|
||||
|
||||
## 快速开始
|
||||
|
||||
```bash
|
||||
# 开发服务器
|
||||
python manage.py runserver 0.0.0.0:8000
|
||||
|
||||
# uWSGI 管理
|
||||
uwsgi --ini uwsgi.ini
|
||||
uwsgi --reload uwsgi.pid
|
||||
uwsgi --stop uwsgi.pid
|
||||
|
||||
# API 文档
|
||||
# /api/docs/ - Swagger UI
|
||||
# /api/redoc/ - ReDoc
|
||||
# /api/schema/ - OpenAPI Schema
|
||||
```
|
||||
|
||||
## 部署
|
||||
|
||||
- 服务器:`simon@doorcome.cn`,路径 `/home/simon/myquant/djapi/`
|
||||
- uWSGI 监听 `127.0.0.1:5004`,nginx 反向代理
|
||||
- 虚拟环境:`/opt/miniconda/envs/django`(Python 3.10)
|
||||
- 域名:`api.doorcome.cn`、`echart.doorcome.cn`
|
||||
|
||||
## 架构
|
||||
|
||||
```
|
||||
djapi/
|
||||
├── djapi/ # 项目配置
|
||||
│ ├── settings.py # Django 设置、CORS、DRF
|
||||
│ ├── urls.py # 根路由 + OpenAPI schema
|
||||
│ └── env_loader.py # .env 加载器
|
||||
├── api/ # 唯一 app
|
||||
│ ├── views.py # 视图层(薄转发)
|
||||
│ ├── urls.py # /api/* 路由
|
||||
│ ├── serializers.py # DRF Serializer(13 个)
|
||||
│ ├── stock/ # 股票数据模块
|
||||
│ │ ├── data_source.py # 统一数据入口(全局单例)
|
||||
│ │ ├── stock_utils.py # 通用工具
|
||||
│ │ ├── stock_basic.py # 日线行情、基本信息
|
||||
│ │ ├── getStockParam.py # 个股参数
|
||||
│ │ ├── getStockEp.py # TTM / 季度 EPS
|
||||
│ │ ├── getIndexs.py # 指数行情
|
||||
│ │ ├── stockMargin.py # 融资融券
|
||||
│ │ ├── getStockFina.py # 财务报表分析
|
||||
│ │ ├── getStockDiv2.py # 股息率计算
|
||||
│ │ └── xwlbDaily.py # 新闻联播数据
|
||||
│ ├── video/ # 新闻联播视频处理(独立模块)
|
||||
│ └── report/ # 日报查询 API
|
||||
│ ├── query.py # 数据库查询
|
||||
│ ├── views.py # 2 个视图
|
||||
│ └── serializers.py # OpenAPI 文档
|
||||
└── uwsgi.ini # uWSGI 配置
|
||||
```
|
||||
|
||||
## 所有 API 端点(16 个)
|
||||
|
||||
基础 URL:`/api/`
|
||||
|
||||
| 端点 | 参数 | 说明 |
|
||||
|------|------|------|
|
||||
| `stockbasic/` | tscode, start_date, end_date | 日线行情 |
|
||||
| `stockinfo/` | tscode | 个股基本信息 |
|
||||
| `stockparam/` | tscode, start_date, end_date | 个股参数(市值等) |
|
||||
| `industrys/` | industry | 按行业查股票列表 |
|
||||
| `indexByName/` | index_name | 按名称查指数 |
|
||||
| `indexDatas/` | tscode, start_date, end_date | 指数日行情 |
|
||||
| `stockep/` | tscode, start_date, end_date | TTM EPS |
|
||||
| `quarterlyEps/` | tscode, start_date, end_date | 季度 EPS |
|
||||
| `finance/` | tscode, start_date, end_date | 财务报表分析 |
|
||||
| `getdiv/` | tscode, start_date, end_date | 股息率(含 TTM) |
|
||||
| `dailymargin/` | trade_date, exchange_id | 每日融资融券汇总 |
|
||||
| `stockmargin/` | tscode, start_date, end_date | 个股融资融券 |
|
||||
| `xwlbNews/` | start_date, end_date | 新闻联播(原始文本) |
|
||||
| `xwlbFine/` | start_date, end_date | 新闻联播(AI 分割后) |
|
||||
| `news/reports/` | report_type, start_date, end_date, id | 日报查询 |
|
||||
| `news/events/` | days, importance, report_type, section, limit | 重要事件聚合 |
|
||||
|
||||
### 数据源
|
||||
|
||||
所有股票数据端点通过 `api/stock/data_source.py` 统一入口:
|
||||
- `get_tushare_pro()` — 全局单例(线程安全)
|
||||
- `get_daily()` — 双源 fallback (Tushare → AkShare)
|
||||
- `get_mysql_db()` — MySQL 全局单例
|
||||
|
||||
## 日报查询 API
|
||||
|
||||
### `GET /api/news/reports/`
|
||||
|
||||
| 参数 | 类型 | 默认 | 说明 |
|
||||
| --- | --- | --- | --- |
|
||||
| `report_type` | string | 两者 | `finance`(A 股)/ `intl`(国际) |
|
||||
| `start_date` | string | 24h 前 | `YYYY-MM-DD` |
|
||||
| `end_date` | string | 今天 | `YYYY-MM-DD` |
|
||||
| `id` | int | 无 | 指定 id 返回单份详情(含事件) |
|
||||
|
||||
### `GET /api/news/events/`
|
||||
|
||||
| 参数 | 类型 | 默认 | 范围 | 说明 |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| `days` | int | 7 | 1~365 | 最近 N 天 |
|
||||
| `importance` | int | 4 | 1~5 | 最低重要度 |
|
||||
| `report_type` | string | 两者 | `finance`/`intl` | 日报类型过滤 |
|
||||
| `section` | string | 全部 | `xwlb`/`news`/`cninfo`/`intl` | 板块过滤 |
|
||||
| `limit` | int | 100 | 1~500 | 返回条数上限 |
|
||||
|
||||
详细说明见 [news_report_api.md](news_report_api.md)。
|
||||
|
||||
## 新闻联播视频处理
|
||||
|
||||
离线批处理流水线:抓取视频 → 下载 → 提取音频 → ASR 转文字 → AI 分割+取标题 → 入库。
|
||||
|
||||
```bash
|
||||
python api/video/main.py # 定时任务入口
|
||||
python api/video/main_videos.py # 批量补缺
|
||||
```
|
||||
|
||||
### 模块文件
|
||||
|
||||
| 文件 | 职责 |
|
||||
|------|------|
|
||||
| `getVideo5.py` | 主流程:抓取→下载→ASR→入库 |
|
||||
| `audioRead.py` | 音频转换、分割、ASR 识别、文本纠错 |
|
||||
| `deepseek.py` | DeepSeek API 封装 |
|
||||
| `newsProcess.py` | AI 新闻分割+标题提取 |
|
||||
| `newsRedo.py` | 手动重处理 |
|
||||
| `main.py` | 定时任务入口 |
|
||||
|
||||
## 部署
|
||||
|
||||
详见 [部署说明](deployment.md)。
|
||||
|
||||
## 文档索引
|
||||
|
||||
| 文档 | 内容 |
|
||||
|------|------|
|
||||
| [使用指南](usage.md) | 各模块使用方法和代码示例 |
|
||||
| [架构说明](architecture.md) | 项目架构、数据流、设计原则 |
|
||||
| [开发指南](development.md) | 环境搭建、开发约定 |
|
||||
| [部署说明](deployment.md) | 本地环境、服务器、uWSGI、rsync 部署 |
|
||||
| [日报查询 API](news_report_api.md) | news/reports + news/events 详细说明 |
|
||||
| [日报数据库](db_schema_v1.1.md) | news_report / news_event 表结构 |
|
||||
@@ -0,0 +1,196 @@
|
||||
# cc-cursor 项目架构
|
||||
|
||||
## 概述
|
||||
|
||||
cc-cursor 是一个 Mac Mini 单机量化研究平台,覆盖从数据获取到策略报告的全链路量化研究流程。
|
||||
|
||||
## 项目布局
|
||||
|
||||
```
|
||||
cc-cursor/ # 项目根目录
|
||||
│
|
||||
├── finance/ # 🔥 核心量化引擎(代码实际位置)
|
||||
│ ├── config/ # 全局配置:.env 加载、数据库/API 路径
|
||||
│ ├── database/ # 数据库层:ORM 模型、SQLAlchemy 连接、DAO
|
||||
│ ├── data/ # 数据层:DataManager 统一入口
|
||||
│ │ └── sources/ # Tushare + AkShare 双数据源实现
|
||||
│ ├── factors/ # 因子引擎:34 个注册因子
|
||||
│ │ ├── technical/ # 技术因子(动量、RSI、MACD、布林等 10 类)
|
||||
│ │ ├── fundamental/ # 基本面因子(ROE、PE、PB、EP)
|
||||
│ │ └── sentiment/ # 情绪因子(Qwen NLP + 三源新闻聚合)
|
||||
│ ├── backtest/ # 回测引擎:VectorBT 封装
|
||||
│ │ ├── vectorbt/ # VectorBT 引擎适配
|
||||
│ │ └── strategies/ # 5 个内置策略(均线、RSI、动量等)
|
||||
│ ├── optimizer/ # 参数优化:Optuna 引擎 + Walk-Forward
|
||||
│ ├── models/ # ML 模型:LightGBM + CatBoost + 特征工程
|
||||
│ │ ├── lightgbm/ # LightGBM 模型封装
|
||||
│ │ └── catboost/ # CatBoost 模型封装
|
||||
│ ├── agents/ # Agent 系统:4 个 Agent + 编排器
|
||||
│ ├── cli/ # 命令行入口:agent_cli + 7 个 demo 验证脚本
|
||||
│ ├── reports/ # 日报输出:daily_YYYYMMDD.md + 存储层
|
||||
│ ├── strategy/ # 策略层(空壳占位,仅 __init__.py)
|
||||
│ ├── portfolio/ # 组合管理(空壳占位,仅 __init__.py)
|
||||
│ ├── execution/ # 执行层(空壳占位,仅 __init__.py)
|
||||
│ └── scheduler/ # 调度层(空壳占位,仅 __init__.py)
|
||||
│
|
||||
├── djapi/ # 🌐 Django API 后端
|
||||
│ ├── djapi/ # 项目配置:settings、urls、env_loader
|
||||
│ └── api/ # 唯一 Django app
|
||||
│ ├── stock/ # A 股数据 API(16 端点)
|
||||
│ ├── video/ # 新闻联播视频处理(独立模块)
|
||||
│ └── report/ # 日报查询 API(news/reports + news/events)
|
||||
│
|
||||
├── shared/ # 🔧 共享工具
|
||||
│ └── script/ # autossh.sh(MariaDB SSH 隧道)
|
||||
│
|
||||
├── docs/ # 📚 项目文档(13 个 md 文件)
|
||||
│
|
||||
├── .claude/ # Claude 配置(空目录,仅保留框架)
|
||||
│
|
||||
├── .git/ # Git 仓库
|
||||
│
|
||||
├── README.md # 项目入口文档
|
||||
└── .gitignore # Git 忽略规则
|
||||
```
|
||||
|
||||
### 各目录职责
|
||||
|
||||
| 目录 | 职责 | 状态 |
|
||||
|------|------|------|
|
||||
| `finance/` | **核心量化引擎**,全链路代码所在地 | ✅ 已实现 |
|
||||
| `finance/config/` | 全局配置,`.env` 加载和路径管理 | ✅ |
|
||||
| `finance/database/` | MariaDB 连接、ORM 模型、DAO 数据访问 | ✅ |
|
||||
| `finance/data/` | 统一数据层,双源(Tushare→AkShare)fallback | ✅ |
|
||||
| `finance/factors/` | 因子引擎,34 因子/12 分类(技术+基本面+情绪) | ✅ |
|
||||
| `finance/backtest/` | VectorBT 回测,5 策略 + 截面回测 | ✅ |
|
||||
| `finance/optimizer/` | Optuna 参数寻优 + Walk-Forward 验证 | ✅ |
|
||||
| `finance/models/` | LightGBM/CatBoost ML 模型 + 特征工程 | ✅ |
|
||||
| `finance/agents/` | 4 Agent(Research/Selection/Risk/Report)+ 编排器 | ✅ |
|
||||
| `finance/cli/` | 命令行入口 + 7 个 demo 验证脚本 | ✅ |
|
||||
| `finance/reports/` | 日报输出(daily_YYYYMMDD.md)+ 持久化存储 | ✅ |
|
||||
| `finance/strategy/` | 策略层,预留扩展 | 🚧 空壳 |
|
||||
| `finance/portfolio/` | 组合管理,预留扩展 | 🚧 空壳 |
|
||||
| `finance/execution/` | 执行层,预留扩展 | 🚧 空壳 |
|
||||
| `finance/scheduler/` | 调度层,预留扩展 | 🚧 空壳 |
|
||||
| `djapi/` | Django API 后端,A 股数据 + 新闻联播 + 日报查询 | ✅ 已部署 |
|
||||
| `shared/` | 跨项目共享工具(SSH 隧道脚本) | ✅ |
|
||||
| `docs/` | 项目文档,13 个 md 文件 | ✅ |
|
||||
|
||||
## 数据流
|
||||
|
||||
```
|
||||
Agent 编排层
|
||||
├── ResearchAgent ── 因子发现(IC/IC_IR 评估)
|
||||
├── SelectionAgent ─ 多因子打分 + ML 预测
|
||||
├── RiskAgent ────── 仓位控制 + 风险预警
|
||||
└── ReportAgent ──── 自动日报生成
|
||||
|
||||
基础引擎层
|
||||
DataManager ──→ FactorEngine ──→ BaseStrategy ──→ VectorBTEngine ──→ BacktestReport
|
||||
│ │ │
|
||||
│ FeatureEngine OptunaEngine
|
||||
│ │ │
|
||||
└──────→ LightGBM/CatBoost ←────────┘
|
||||
|
||||
情绪增强层
|
||||
NewsSource(AkShare/DB/MCP) ──→ QwenClient ──→ SentimentFactor ──→ FactorEngine
|
||||
```
|
||||
|
||||
### 核心数据流
|
||||
|
||||
```
|
||||
Data → Factor → Model → Strategy → Backtest → Report
|
||||
```
|
||||
|
||||
## 设计原则
|
||||
|
||||
### 模块隔离
|
||||
各引擎通过统一接口交互,可替换实现(VectorBT → Backtrader)。
|
||||
|
||||
### 接口标准化
|
||||
- 因子:`calculate(df) → pd.Series`
|
||||
- 策略:`generate_signals(df) → pd.Series`
|
||||
- 模型:`fit/predict/save/load`
|
||||
- 优化:`optimize() → OptimizationResult`
|
||||
|
||||
### 数据层统一
|
||||
策略/模型不直连数据源,全部通过 `DataManager`。禁止:
|
||||
- 策略直接访问 AkShare/Tushare
|
||||
- 模型直接访问数据库
|
||||
|
||||
### Agent 不重建轮子
|
||||
Agent 通过依赖注入复用已有引擎,编排而非重建。
|
||||
|
||||
### 防前视偏差
|
||||
- 时间序列交叉验证(TimeSeriesSplit)
|
||||
- expanding window 统计量
|
||||
- 特征工程 fit 在训练集,transform 在测试集
|
||||
|
||||
## 技术栈
|
||||
|
||||
| 组件 | 技术 | 说明 |
|
||||
|------|------|------|
|
||||
| 数据获取 | AkShare + Tushare (双源) | Tushare 优先,AkShare fallback |
|
||||
| 数据库 | MariaDB (SSH 隧道) | 本地 13306 → 远程 3306 |
|
||||
| 因子/特征 | pandas / numpy / sklearn | — |
|
||||
| 回测引擎 | VectorBT 1.0 | 只做多,10 万/万三 |
|
||||
| 参数优化 | Optuna 4.9 | Walk-Forward 验证 |
|
||||
| ML 模型 | LightGBM 4.6 + CatBoost 1.2 | 统一接口 |
|
||||
| NLP 情绪 | Qwen (DashScope / Ollama) | 双后端 |
|
||||
| Agent 编排 | 自研编排器 | finance/agents/ |
|
||||
| API 后端 | Django 5.2 + uWSGI | djapi/ |
|
||||
|
||||
## 开发进度
|
||||
|
||||
| Sprint | 模块 | 状态 |
|
||||
|--------|------|------|
|
||||
| Sprint 0 | 基础设施(DataManager + MariaDB 3 表) | ✅ |
|
||||
| Sprint 1 | 因子引擎(34 因子 / 12 分类) | ✅ |
|
||||
| Sprint 2 | VectorBT 回测(5 策略 + 截面 + BacktestReport) | ✅ |
|
||||
| Sprint 3 | Optuna 优化(+ Walk-Forward) | ✅ |
|
||||
| Sprint 4 | ML 模型(LightGBM + CatBoost + MLStrategy) | ✅ |
|
||||
| Sprint 5 | Qwen 情绪因子(三源新闻 + 日期对齐) | ✅ |
|
||||
| Sprint 6 | Agent 系统(4 Agent + CLI + 日报 .md/.html) | ✅ |
|
||||
| Sprint 7 | djapi API(日报查询 ×2) | ✅ |
|
||||
|
||||
**全部 8 个 Sprint 已完成。**
|
||||
|
||||
## 数据源架构
|
||||
|
||||
### finance/ 引擎层
|
||||
```
|
||||
DataManager
|
||||
├── TushareSource(优先,需 TUSHARE_TOKEN)
|
||||
├── AkShareSource(fallback,无需 token)
|
||||
└── Database Cache(SQLAlchemy + MariaDB)
|
||||
```
|
||||
|
||||
### djapi/ API 层
|
||||
```
|
||||
api/stock/data_source.py(统一入口)
|
||||
├── get_tushare_pro() — 全局单例
|
||||
├── get_daily() — 双源 fallback
|
||||
└── get_mysql_db() — MySQL 全局单例
|
||||
```
|
||||
|
||||
## 指数代码规则
|
||||
|
||||
- `.SH` 结尾且不以 `399` 开头 → 指数(如 `000001.SH`)
|
||||
- `.SZ` 开头非 `399` → 个股(如 `000001.SZ`)
|
||||
- `399*.SZ` → 指数(如 `399001.SZ`)
|
||||
|
||||
## 文档索引
|
||||
|
||||
| 文档 | 内容 |
|
||||
|------|------|
|
||||
| [使用指南](usage.md) | 各模块使用方法和代码示例 |
|
||||
| [开发指南](development.md) | 环境搭建、开发约定、模块说明 |
|
||||
| [因子与表结构速查](reference.md) | 34 因子注册表、DB 表结构、数据源接口 |
|
||||
| [数据层详解](data-layer.md) | DataManager、数据库、缓存策略、已知 Bug |
|
||||
| [因子引擎详解](factors.md) | 因子计算、情绪引擎、新闻源 |
|
||||
| [回测引擎详解](backtest.md) | VectorBT、策略、信号工具、Optuna |
|
||||
| [ML 模型详解](ml-models.md) | 特征工程、LightGBM/CatBoost、ML 策略 |
|
||||
| [Agent 系统详解](agents.md) | Agent 架构、CLI、日报 |
|
||||
| [部署说明](deployment.md) | 本地环境、服务器、uWSGI、rsync 部署 |
|
||||
| [DJAPI 接口](api.md) | Django API 端点参考 |
|
||||
| [日报查询 API](news_report_api.md) | news/reports + news/events 接口 |
|
||||
@@ -1,4 +1,4 @@
|
||||
# CLAUDE-backtest.md — 回测引擎 + 参数优化
|
||||
# 回测引擎 + 参数优化
|
||||
|
||||
## VectorBTEngine (`finance/backtest/vectorbt/engine.py`)
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# CLAUDE-data.md — 数据层 + 数据库
|
||||
# 数据层 + 数据库
|
||||
|
||||
## DataManager (`finance/data/data_manager.py`)
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
# 日报结构化入库:数据库表结构与数据契约
|
||||
|
||||
> 版本:v1.0 | 2026-08-03
|
||||
> 版本:v1.1 | 2026-08-06
|
||||
> 变更(v1.0 → v1.1):新增 §3.1 stats 口径说明,修正 §3 中 pipeline "M1→M6" 的过时描述(实际 key 为 raw_total/raw_by_source/proc 等)。
|
||||
> 用途:供 API / 前端对接读取日报数据。表位于 MySQL `myquant` 库,表前缀 `news_`。
|
||||
> 连接:`192.168.1.10:13306`(pi 上 autossh 隧道 → doorcome.cn:3306 MariaDB 10.11),用户 `myquant`(密码在服务器 `.env` 的 `NEWS_DB_PASSWORD`)。
|
||||
|
||||
@@ -37,6 +38,7 @@
|
||||
| summary | TEXT NULL | 摘要/正文 |
|
||||
| sentiment | VARCHAR(8) NULL | `positive` / `negative` / `neutral` |
|
||||
| source | VARCHAR(64) NULL | 来源(如 `cls`、`investinglive.com`) |
|
||||
| sources | TEXT NULL | **该新闻全部来源**,JSON 数组字符串(如 `["yicai","stcn"]`);无多源/历史数据可为 `null` |
|
||||
| url | VARCHAR(512) NULL | 原文链接(新闻联播为空) |
|
||||
| created_at | DATETIME | 入库时间 |
|
||||
|
||||
@@ -59,7 +61,7 @@
|
||||
|
||||
| key | finance | intl | 内容 |
|
||||
| --- | --- | --- | --- |
|
||||
| `pipeline` | ✅ | ✅ | M1→M6 管道各环节数量:`{label: 数量}` |
|
||||
| `pipeline` | ✅ | ✅ | 管道各环节数量。**finance 与 intl 内部 key 集不同**:finance 为 `{raw_total, raw_by_source, raw_total_24h, raw_by_source_24h, proc, deduped, dups, emb_count, qdrant_count, cninfo_raw}`;intl 为 `{raw_total, processed, deduped, embedded, qdrant}`(口径见 §3.1) |
|
||||
| `sources` | ✅ | — | 各新闻源文章数:`{源名: 数量}` |
|
||||
| `news` | ✅ | — | 新闻统计:`{total, hi_threshold, sentiments, importances, event_types}` |
|
||||
| `cninfo` | ✅ | — | 公告调研统计:`{total, hi_threshold, by_day, announcement, research, irm}` |
|
||||
@@ -71,6 +73,22 @@
|
||||
|
||||
> 历史文件与新生成日报的 stats 结构存在差异(历史为 HTML 解析快照,新生成为结构化组装),前端建议按 key 防御性读取。
|
||||
|
||||
### 3.1 口径说明(重要,避免误解)
|
||||
|
||||
`stats` 内各数字口径不同,请勿直接互相比较:
|
||||
|
||||
| 字段 | 口径 |
|
||||
| --- | --- |
|
||||
| `pipeline.raw_total` | **日报日期当天**抓取的文章数(`data/raw/{src}/{date}/index.jsonl` 中 `stage=article 且 success` 的条目)。`raw_by_source` 是各源明细,**其和 = raw_total**;当天未抓取/无文章的源显示 0 |
|
||||
| `pipeline.raw_total_24h` | 最近 24 小时内**抓取**(按 `fetched_at`)的文章数;`raw_by_source_24h` 为各源明细,和 = raw_total_24h。当天 07:00 抓取的数据其值 ≈ raw_total(并非"24h 内发布的新闻",raw 层无发布时间的可靠字段) |
|
||||
| `pipeline.proc / deduped / dups / emb_count / qdrant_count` | 抽取 / 去重后 / 重复 / 向量化 / Qdrant 总量(`qdrant_count` 为全量累计,非当天) |
|
||||
| `news.total` | **过去 30 小时窗口内**经 LLM 抽取的新闻事件数。**≠ raw_total**:raw 是抓取的文章数,news 是抽取后的事件数(会有过滤/合并),两者不可互相验证 |
|
||||
| `news.importances` | `{重要度等级(1-5): 事件数}`,**各等级之和 = news.total** |
|
||||
| `news.sentiments` | `{情绪: 事件数}`(positive/negative/neutral),和 = news.total |
|
||||
| `news.event_types` | `{事件类型: 事件数}`(TOP 10) |
|
||||
| `pipeline.*(finance)` | finance 日报专用:`proc`=抽取后条数、`deduped`=去重后、`dups`=重复条数、`emb_count`=向量化条数、`qdrant_count`=Qdrant 全量累计(非当天)、`cninfo_raw`=当天 cninfo 抓取数。`raw_by_source` 各源之和 = `raw_total` |
|
||||
| `pipeline.*(intl)` | intl 日报**结构不同**:`raw_total`=当天抓取文章数、`processed`=处理数(≈ raw_total)、`deduped`=去重后条数、`embedded`=向量化条数、`qdrant`=Qdrant 全量累计。intl 无 `raw_by_source` 明细与 `cninfo_raw` |
|
||||
|
||||
---
|
||||
|
||||
## 4. 常用查询示例(API 实现参考)
|
||||
@@ -0,0 +1,274 @@
|
||||
# 部署说明
|
||||
|
||||
cc-cursor 包含两个运行组件:**finance 量化引擎**(Mac Mini 本地)和 **djapi API 后端**(Linux 服务器)。部署架构如下:
|
||||
|
||||
```
|
||||
┌─ Mac Mini (本地) ─────────────────────────────────────────────┐
|
||||
│ │
|
||||
│ finance/ 量化引擎 │
|
||||
│ ├── 数据获取 (AkShare/Tushare) │
|
||||
│ ├── 因子计算 / 回测 / ML │
|
||||
│ └── Agent 日报生成 │
|
||||
│ │
|
||||
│ shared/script/autossh.sh │
|
||||
│ └── SSH 隧道 :13306 ──────────────────────┐ │
|
||||
│ │ │
|
||||
└─────────────────────────────────────────────┼──────────────────┘
|
||||
│
|
||||
MariaDB 10.11 │
|
||||
doorcome.cn:3306│
|
||||
│
|
||||
┌─ Linux 服务器 (doorcome.cn) ────────────────┼──────────────────┐
|
||||
│ │ │
|
||||
│ djapi/ Django API │ │
|
||||
│ ├── uWSGI :5004 │ │
|
||||
│ ├── nginx 反向代理 │ │
|
||||
│ └── 域名: api.doorcome.cn │ │
|
||||
│ │ │
|
||||
│ MariaDB myquant 库 │ │
|
||||
│ ├── mac_* 表 (量化引擎数据) │ │
|
||||
│ ├── xwlb_* 表 (新闻联播) │ │
|
||||
│ └── news_* 表 (日报) │ │
|
||||
│ │ │
|
||||
└─────────────────────────────────────────────┘ │
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 1. 本地环境(Mac Mini)
|
||||
|
||||
### 1.1 Python 环境
|
||||
|
||||
```bash
|
||||
conda activate quant # Python 3.11.13
|
||||
```
|
||||
|
||||
### 1.2 环境变量
|
||||
|
||||
复制并编辑 `finance/.env`(参考 `finance/.env.example`):
|
||||
|
||||
```bash
|
||||
TUSHARE_TOKEN=your_token_here
|
||||
QWEN_API_KEY=sk-your-key-here
|
||||
MAC_DB_HOST=127.0.0.1
|
||||
MAC_DB_PORT=13306
|
||||
MAC_DB_USER=myquant
|
||||
MAC_DB_PASSWORD=your_password_here
|
||||
MAC_DB_NAME=myquant
|
||||
```
|
||||
|
||||
### 1.3 数据库 SSH 隧道
|
||||
|
||||
```bash
|
||||
bash shared/script/autossh.sh
|
||||
```
|
||||
|
||||
脚本内容:
|
||||
|
||||
```bash
|
||||
autossh -M 0 -fN -L 13306:localhost:3306 tunnel@doorcome.cn
|
||||
```
|
||||
|
||||
验证隧道:
|
||||
|
||||
```bash
|
||||
lsof -i :13306 | grep LISTEN
|
||||
```
|
||||
|
||||
连接信息:
|
||||
|
||||
```
|
||||
Host: 127.0.0.1
|
||||
Port: 13306
|
||||
User: myquant
|
||||
Database: myquant
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. 服务器环境(doorcome.cn)
|
||||
|
||||
### 2.1 基本信息
|
||||
|
||||
| 项目 | 值 |
|
||||
|------|-----|
|
||||
| 服务器 | `simon@doorcome.cn` |
|
||||
| 项目路径 | `/home/simon/myquant/djapi/` |
|
||||
| Python 环境 | `/opt/miniconda/envs/django` (Python 3.10) |
|
||||
| uWSGI 端口 | `127.0.0.1:5004` |
|
||||
| nginx 反向代理 | 域名 → `127.0.0.1:5004` |
|
||||
| 生产域名 | `api.doorcome.cn`、`echart.doorcome.cn` |
|
||||
|
||||
### 2.2 服务器环境变量
|
||||
|
||||
服务器端 `djapi/.env`(**不随代码同步**,需在服务器上手动维护):
|
||||
|
||||
```bash
|
||||
DJANGO_SECRET_KEY=...
|
||||
TUSHARE_TS_TOKEN=...
|
||||
MYSQL_HOST=localhost
|
||||
MYSQL_PORT=3306
|
||||
MYSQL_USER=myquant
|
||||
MYSQL_PASSWORD=...
|
||||
MYSQL_DATABASE=myquant
|
||||
NEWS_DB_HOST=127.0.0.1
|
||||
NEWS_DB_PORT=3306
|
||||
NEWS_DB_USER=myquant
|
||||
NEWS_DB_PASSWORD=...
|
||||
NEWS_DB_NAME=myquant
|
||||
DEEPSEEK_API_KEY=...
|
||||
DASHSCOPE_API_KEY=...
|
||||
```
|
||||
|
||||
### 2.3 uWSGI 配置
|
||||
|
||||
配置文件:`djapi/uwsgi.ini`
|
||||
|
||||
```ini
|
||||
[uwsgi]
|
||||
http = 127.0.0.1:5004
|
||||
chdir = /home/simon/myquant/djapi
|
||||
module = djapi.wsgi:application
|
||||
uid = simon
|
||||
gid = simon
|
||||
master = true
|
||||
workers = 5
|
||||
pidfile = /home/simon/myquant/djapi/uwsgi.pid
|
||||
vacuum = true
|
||||
thunder-lock = true
|
||||
enable-threads = true
|
||||
harakiri = 30
|
||||
post-buffering = 4096
|
||||
daemonize = /home/simon/myquant/djapi/uwsgi.log
|
||||
log-maxsize = 10240000
|
||||
py-autoreload = 1
|
||||
virtualenv = /opt/miniconda/envs/django
|
||||
env = DJANGO_SETTINGS_MODULE=djapi.settings
|
||||
```
|
||||
|
||||
### 2.4 uWSGI 管理命令
|
||||
|
||||
```bash
|
||||
# 启动
|
||||
uwsgi --ini uwsgi.ini
|
||||
|
||||
# 热重载
|
||||
uwsgi --reload uwsgi.pid
|
||||
|
||||
# 停止
|
||||
uwsgi --stop uwsgi.pid
|
||||
|
||||
# 强制停止
|
||||
kill $(lsof -ti:5004)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. 代码部署
|
||||
|
||||
### 3.1 全量同步
|
||||
|
||||
```bash
|
||||
rsync -avz --delete \
|
||||
--exclude='.env' --exclude='db.sqlite3' \
|
||||
--exclude='*.log' --exclude='uwsgi.pid' \
|
||||
--exclude='__pycache__/' --exclude='*.pyc' \
|
||||
--exclude='xwlb_video/' --exclude='audio_processing/' \
|
||||
/path/to/cc-cursor/djapi/ \
|
||||
simon@doorcome.cn:/home/simon/myquant/djapi/
|
||||
```
|
||||
|
||||
### 3.2 单文件同步
|
||||
|
||||
```bash
|
||||
# 必须写完整目标路径,否则会展平到根目录
|
||||
rsync -avz api/views.py simon@doorcome.cn:/home/simon/myquant/djapi/api/views.py
|
||||
```
|
||||
|
||||
### 3.3 重启服务
|
||||
|
||||
```bash
|
||||
ssh simon@doorcome.cn "kill \$(lsof -ti:5004); sleep 2; /opt/miniconda/envs/django/bin/uwsgi --ini /home/simon/myquant/djapi/uwsgi.ini"
|
||||
```
|
||||
|
||||
### 3.4 部署后验证
|
||||
|
||||
```bash
|
||||
# Swagger 文档
|
||||
curl -s https://api.doorcome.cn/api/docs/ | head -5
|
||||
|
||||
# 日报查询
|
||||
curl -s 'https://api.doorcome.cn/api/news/reports/' | python -m json.tool | head -20
|
||||
|
||||
# 健康检查(冒烟)
|
||||
curl -s -o /dev/null -w "%{http_code}" 'https://api.doorcome.cn/api/news/reports/'
|
||||
# → 200
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. 数据库
|
||||
|
||||
### 4.1 表前缀
|
||||
|
||||
| 前缀 | 用途 | 位置 |
|
||||
|------|------|------|
|
||||
| `mac_` | 量化引擎数据(股票列表、日线、财务、报告) | finance 引擎写入 |
|
||||
| `xwlb_` | 新闻联播数据 | djapi video 模块写入 |
|
||||
| `news_` | 日报数据 | 外部 pipeline 写入,djapi 只读 |
|
||||
|
||||
### 4.2 量化引擎表
|
||||
|
||||
| 表 | 内容 | 主键 |
|
||||
|----|------|------|
|
||||
| `mac_stock_basic` | A 股列表 (5,524 只) | ts_code |
|
||||
| `mac_stock_daily` | 日线 OHLCV | (ts_code, trade_date) |
|
||||
| `mac_stock_financial` | 财务指标 | (ts_code, end_date) |
|
||||
| `mac_report` | 报告持久化 | id |
|
||||
|
||||
### 4.3 日报表
|
||||
|
||||
| 表 | 内容 |
|
||||
|----|------|
|
||||
| `news_report` | 日报主表(一行 = 一份日报) |
|
||||
| `news_event` | 日报事件明细(一行 = 一条事件) |
|
||||
|
||||
详见 [db_schema_v1.1.md](db_schema_v1.1.md)。
|
||||
|
||||
---
|
||||
|
||||
## 5. 开发环境
|
||||
|
||||
### 5.1 本地运行 djapi
|
||||
|
||||
```bash
|
||||
cd djapi
|
||||
python manage.py runserver 0.0.0.0:8000
|
||||
```
|
||||
|
||||
### 5.2 API 文档
|
||||
|
||||
- Swagger UI:`/api/docs/`
|
||||
- ReDoc:`/api/redoc/`
|
||||
- OpenAPI Schema:`/api/schema/`
|
||||
|
||||
### 5.3 数据库初始化
|
||||
|
||||
```bash
|
||||
python manage.py makemigrations
|
||||
python manage.py migrate
|
||||
python manage.py createsuperuser
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. 文档索引
|
||||
|
||||
| 文档 | 内容 |
|
||||
|------|------|
|
||||
| [使用指南](usage.md) | 各模块使用方法和代码示例 |
|
||||
| [架构说明](architecture.md) | 项目架构、数据流、设计原则 |
|
||||
| [开发指南](development.md) | 环境搭建、开发约定 |
|
||||
| [DJAPI 接口](api.md) | Django API 端点参考 |
|
||||
| [日报查询 API](news_report_api.md) | news/reports + news/events 接口 |
|
||||
| [日报数据库](db_schema_v1.1.md) | news_report / news_event 表结构 |
|
||||
@@ -0,0 +1,271 @@
|
||||
# cc-cursor 开发指南
|
||||
|
||||
## 环境搭建
|
||||
|
||||
### 1. 克隆项目
|
||||
|
||||
```bash
|
||||
git clone https://github.com/Simon2046/myquant.git
|
||||
cd cc-cursor
|
||||
```
|
||||
|
||||
### 2. Python 环境
|
||||
|
||||
```bash
|
||||
conda activate quant # Python 3.11.13
|
||||
```
|
||||
|
||||
### 3. 数据库 SSH 隧道
|
||||
|
||||
```bash
|
||||
bash shared/script/autossh.sh
|
||||
# host: 127.0.0.1:13306 user: myquant database: myquant
|
||||
```
|
||||
|
||||
### 4. 环境变量配置
|
||||
|
||||
复制并编辑 `finance/.env`(参考 `finance/.env.example`):
|
||||
|
||||
```bash
|
||||
TUSHARE_TOKEN=your_token_here
|
||||
QWEN_API_KEY=sk-your-key-here
|
||||
MAC_DB_PASSWORD=your_password_here
|
||||
```
|
||||
|
||||
### 5. 验证环境
|
||||
|
||||
```bash
|
||||
python finance/cli/demo_data_manager.py
|
||||
```
|
||||
|
||||
## 项目结构
|
||||
|
||||
```
|
||||
finance/ # 核心量化引擎(代码实际位置)
|
||||
├── config/ # 全局配置
|
||||
│ └── settings.py # .env 加载 + 路径配置
|
||||
├── database/ # 数据库层
|
||||
│ ├── connection.py # SQLAlchemy 引擎 + SSH 自动恢复
|
||||
│ ├── models.py # ORM 模型(mac_ 前缀表)
|
||||
│ └── dao.py # 数据访问对象
|
||||
├── data/ # 数据层
|
||||
│ ├── data_manager.py # 统一数据入口
|
||||
│ └── sources/ # 数据源实现
|
||||
│ ├── tushare_source.py
|
||||
│ └── akshare_source.py
|
||||
├── factors/ # 因子引擎
|
||||
│ ├── base.py # 因子基类
|
||||
│ ├── engine.py # 因子计算引擎
|
||||
│ ├── registry.py # 因子注册表
|
||||
│ ├── technical/ # 技术因子(10 类)
|
||||
│ ├── fundamental/ # 基本面因子
|
||||
│ └── sentiment/ # 情绪因子
|
||||
│ ├── sentiment_engine.py
|
||||
│ ├── sentiment_factor.py
|
||||
│ ├── news_source.py # 三源新闻聚合
|
||||
│ └── qwen_client.py # Qwen API 客户端
|
||||
├── backtest/ # 回测引擎
|
||||
│ ├── base.py # 策略基类
|
||||
│ ├── report.py # 回测报告
|
||||
│ ├── signal.py # 信号工具
|
||||
│ ├── vectorbt/ # VectorBT 引擎
|
||||
│ └── strategies/ # 内置策略
|
||||
├── optimizer/ # 参数优化
|
||||
│ ├── engine.py # Optuna 引擎
|
||||
│ ├── space.py # 搜索空间
|
||||
│ ├── objectives.py # 优化目标
|
||||
│ └── result.py # 优化结果
|
||||
├── models/ # ML 模型
|
||||
│ ├── base.py # 模型基类
|
||||
│ ├── features.py # 特征工程
|
||||
│ ├── backtest_integration.py # ML 策略
|
||||
│ ├── lightgbm/
|
||||
│ └── catboost/
|
||||
├── agents/ # Agent 系统
|
||||
│ ├── base.py # Agent 基类
|
||||
│ ├── orchestrator.py # 编排器
|
||||
│ ├── research_agent.py # 因子研究
|
||||
│ ├── selection_agent.py # 股票打分
|
||||
│ ├── risk_agent.py # 风险评估
|
||||
│ └── report_agent.py # 日报生成
|
||||
├── cli/ # 命令行
|
||||
│ ├── agent_cli.py # Agent CLI 入口
|
||||
│ └── demo_*.py # 验证脚本
|
||||
└── reports/ # 日报输出
|
||||
└── storage.py # 报告持久化
|
||||
```
|
||||
|
||||
## 开发约定
|
||||
|
||||
### 代码组织
|
||||
|
||||
- 代码内 import 用顶层名 `data.*`/`factors.*` 等 — CLI 自动把 `finance/` 加入 sys.path
|
||||
- 文件路径为 `finance/data/xxx.py` 等
|
||||
|
||||
### 数据流约束
|
||||
|
||||
```
|
||||
Data → Factor → Model → Strategy → Backtest → Report
|
||||
```
|
||||
|
||||
**必须遵守**:
|
||||
- 策略层禁止直接访问 AkShare/Tushare → 全部通过 `DataManager`
|
||||
- 模型层禁止直接访问数据库 → 全部通过 `DataManager`
|
||||
- 指数代码规则:`.SH`=指数, `.SZ` 开头非 399=个股
|
||||
- 数据源优先级:Tushare → AkShare (fallback)
|
||||
|
||||
### 接口规范
|
||||
|
||||
#### 因子接口
|
||||
```python
|
||||
class BaseFactor:
|
||||
def calculate(self, df: pd.DataFrame) -> pd.Series:
|
||||
"""接收 OHLCV 数据,返回因子值序列"""
|
||||
```
|
||||
|
||||
#### 策略接口
|
||||
```python
|
||||
class BaseStrategy:
|
||||
def generate_signals(self, factor_df: pd.DataFrame) -> pd.Series:
|
||||
"""接收因子数据,返回交易信号: 1=buy, 0=sell, -1=hold"""
|
||||
```
|
||||
|
||||
#### 模型接口
|
||||
```python
|
||||
class BaseModel:
|
||||
def fit(self, X, y): ...
|
||||
def predict(self, X) -> np.ndarray: ...
|
||||
def save(self, path): ...
|
||||
@classmethod
|
||||
def load(cls, path): ...
|
||||
```
|
||||
|
||||
### 多步任务规则
|
||||
|
||||
复杂任务(涉及 3+ 文件或 2+ 模块)执行前:
|
||||
1. 输出执行计划清单(步骤 + 每步验证方法)
|
||||
2. 每步完成后验证通过才继续
|
||||
3. 遇到失败先定位根因,不跳过
|
||||
|
||||
### 修改多文件前
|
||||
|
||||
先说明:文件清单、原因、影响;优先小范围修改。
|
||||
|
||||
## 模块说明
|
||||
|
||||
### 数据层 (`finance/data/`)
|
||||
|
||||
双数据源架构:Tushare(优先)→ AkShare(fallback),DB 缓存优先。
|
||||
|
||||
```python
|
||||
from data.data_manager import DataManager
|
||||
dm = DataManager()
|
||||
dm.init_db() # 首次建表(幂等)
|
||||
stocks = dm.get_stock_list() # → 5,524 只
|
||||
daily = dm.get_daily("000001.SZ") # → 日线
|
||||
fina = dm.get_financial("000001.SZ") # → 财务
|
||||
n = dm.sync_daily("000001.SZ") # → 增量同步
|
||||
```
|
||||
|
||||
### 因子引擎 (`finance/factors/`)
|
||||
|
||||
34 个注册因子,12 个分类:动量、RSI、MACD、量价、布林、ATR、均线、波动率、换手率、振幅、基本面、情绪。
|
||||
|
||||
```python
|
||||
from factors.registry import get_factor, list_factors
|
||||
from factors.engine import FactorEngine
|
||||
|
||||
fe = FactorEngine(dm)
|
||||
factor_df = fe.compute("000001.SZ", [get_factor("momentum_20"), get_factor("rsi_14")])
|
||||
```
|
||||
|
||||
### 回测引擎 (`finance/backtest/`)
|
||||
|
||||
VectorBT 1.0,只做多,10 万/万三。5 个内置策略 + 自定义策略接口。
|
||||
|
||||
```python
|
||||
from backtest.vectorbt.engine import VectorBTEngine
|
||||
engine_bt = VectorBTEngine(initial_capital=100_000, commission=0.0003)
|
||||
report = engine_bt.run(strategy, price_df, factor_df)
|
||||
```
|
||||
|
||||
### 参数优化 (`finance/optimizer/`)
|
||||
|
||||
Optuna 4.9 + Walk-Forward 滚动验证。
|
||||
|
||||
```python
|
||||
from optimizer.engine import OptunaEngine
|
||||
opt = OptunaEngine(engine_bt)
|
||||
result = opt.optimize(StrategyClass, space, price_df, factor_df, metric="sharpe", n_trials=200)
|
||||
```
|
||||
|
||||
### ML 模型 (`finance/models/`)
|
||||
|
||||
LightGBM 4.6 + CatBoost 1.2,统一接口,特征工程防前视偏差。
|
||||
|
||||
```python
|
||||
from models.features import FeatureEngine
|
||||
from models.lightgbm.model import LightGBMModel
|
||||
|
||||
fe = FeatureEngine(lookahead=5)
|
||||
X, y = fe.build(factor_df, price_df, fit=True)
|
||||
model = LightGBMModel(params={"n_estimators": 200}).fit(X_train, y_train)
|
||||
```
|
||||
|
||||
### 情绪因子 (`finance/factors/sentiment/`)
|
||||
|
||||
三数据源聚合(AkShare 个股新闻 + 新闻联播 DB + MCP trendradar-news),Qwen DashScope + Ollama 双后端。
|
||||
|
||||
```python
|
||||
from factors.sentiment.sentiment_engine import SentimentEngine
|
||||
sent = SentimentEngine(dm)
|
||||
sent_df = sent.compute("000001.SZ", max_news=20)
|
||||
```
|
||||
|
||||
### Agent 系统 (`finance/agents/`)
|
||||
|
||||
4 个 Agent(Research/Selection/Risk/Report)+ 编排器 + CLI。
|
||||
|
||||
```bash
|
||||
python finance/cli/agent_cli.py daily # 完整流程
|
||||
python finance/cli/agent_cli.py picks 15 # 选股
|
||||
python finance/cli/agent_cli.py risk # 风险评估
|
||||
```
|
||||
|
||||
## 安全规范
|
||||
|
||||
### 禁止提交
|
||||
|
||||
- `.env`(含真实 key)
|
||||
- API Key(`sk-*`、`TUSHARE_TOKEN` 等)
|
||||
- Cookie / Session / Token / 密钥
|
||||
- 个人隐私数据(手机号、身份证、密码)
|
||||
|
||||
### 必须提供
|
||||
|
||||
- `.env.example` — 仅含占位符的示例配置
|
||||
|
||||
### 提交前检查
|
||||
|
||||
```bash
|
||||
grep -r "sk-\|token\|password" --include="*.py" --include="*.md" --include="*.yaml" | grep -v ".example\|your_token\|your_password"
|
||||
```
|
||||
|
||||
## 已知 Bug 速查
|
||||
|
||||
详见 [data-layer.md](data-layer.md) 和 [agents.md](agents.md) 中的 Bug 列表。
|
||||
|
||||
## 文档索引
|
||||
|
||||
| 文档 | 内容 |
|
||||
|------|------|
|
||||
| [使用指南](usage.md) | 各模块使用方法和代码示例 |
|
||||
| [架构说明](architecture.md) | 项目架构、数据流、设计原则 |
|
||||
| [因子与表结构速查](reference.md) | 34 因子注册表、DB 表结构、数据源接口 |
|
||||
| [数据层详解](data-layer.md) | DataManager、数据库、缓存策略、已知 Bug |
|
||||
| [因子引擎详解](factors.md) | 因子计算、情绪引擎、新闻源 |
|
||||
| [回测引擎详解](backtest.md) | VectorBT、策略、信号工具、Optuna |
|
||||
| [ML 模型详解](ml-models.md) | 特征工程、LightGBM/CatBoost、ML 策略 |
|
||||
| [Agent 系统详解](agents.md) | Agent 架构、CLI、日报 |
|
||||
| [部署说明](deployment.md) | 本地环境、服务器、uWSGI、rsync 部署 |
|
||||
| [DJAPI 接口](api.md) | Django API 端点参考 |
|
||||
@@ -1,4 +1,4 @@
|
||||
# CLAUDE-factors.md — 因子引擎 + 情绪因子
|
||||
# 因子引擎 + 情绪因子
|
||||
|
||||
## FactorEngine (`finance/factors/engine.py`)
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# CLAUDE-ml.md — ML 模型层
|
||||
# ML 模型层
|
||||
|
||||
## FeatureEngine (`finance/models/features.py`)
|
||||
|
||||
+35
-13
@@ -1,7 +1,7 @@
|
||||
# 日报查询 API 使用手册(news_report / news_event)
|
||||
|
||||
> 版本:v1.1 | 2026-08-03(已部署至生产 `api.doorcome.cn`)
|
||||
> 数据表定义见 `djapi/docs/db_schema.md`,数据生产侧设计见 `djapi/docs/report_db_design.md`。
|
||||
> 数据表定义见 [db_schema_v1.1.md](db_schema_v1.1.md)。
|
||||
> 本文档为前端/调用方对接手册:两个只读查询接口的线上地址、参数、curl 用法与返回结构。
|
||||
|
||||
---
|
||||
@@ -61,11 +61,22 @@ Swagger 交互式文档(自动生成,含全部参数说明):`https://api
|
||||
"generated_at": "2026-08-03T07:00:00",
|
||||
"ai_summary": "……(AI 摘要全文,按条目分行)",
|
||||
"stats": {
|
||||
"pipeline": { "M1": 210, "M2": 205, "M3": 198, "M4": 190, "M5": 188, "M6": 180 },
|
||||
"sources": { "cls": 95, "eastmoney": 60, "other": 25 },
|
||||
"news": { "total": 180, "hi_threshold": 18, "sentiments": {...}, "importances": [...], "event_types": [...] },
|
||||
"cninfo": { "total": 58, "hi_threshold": 4, "by_day": {...}, "announcement": 40, "research": 15, "irm": 3 },
|
||||
"xwlb": { "total": 28, "date": "2026-08-03" }
|
||||
"pipeline": {
|
||||
"raw_total": 116,
|
||||
"raw_total_24h": 116,
|
||||
"raw_by_source": { "经济观察网": 31, "新浪财经": 27, "证券时报": 28, "财联社": 15, "东方财富": 7, "中证券网": 2, "第一财经": 3, "中国证券网": 3 },
|
||||
"raw_by_source_24h": { ... },
|
||||
"proc": ..., "deduped": ..., "dups": ..., "emb_count": ..., "qdrant_count": ..., "cninfo_raw": ...
|
||||
},
|
||||
"news": {
|
||||
"total": 643,
|
||||
"hi_threshold": 4,
|
||||
"sentiments": { "neutral": 470, "positive": 123, "negative": 50 },
|
||||
"importances": { "1": 115, "2": 275, "3": 185, "4": 48, "5": 20 },
|
||||
"event_types": { "其他": 287, "国际局势": 107, "宏观政策": 62, "行业政策": 49, "财报披露": 29 }
|
||||
},
|
||||
"cninfo": { "total": 48, "hi_threshold": 2, "by_day": { "2026-08-06": 8, "2026-08-01": 9 }, "announcement": 47, "research": 1, "irm": 0 },
|
||||
"xwlb": { "total": 21, "date": "08月05日" }
|
||||
},
|
||||
"created_at": "2026-08-03T07:01:00"
|
||||
},
|
||||
@@ -77,18 +88,26 @@ Swagger 交互式文档(自动生成,含全部参数说明):`https://api
|
||||
"generated_at": "2026-08-03T07:00:00",
|
||||
"ai_summary": "……",
|
||||
"stats": {
|
||||
"pipeline": { "M1": 150, "M2": 145, "M3": 140, "M4": 135, "M5": 130, "M6": 125 },
|
||||
"sentiment": [...],
|
||||
"importance": [{ "重要度": 5, "数量": 3 }, { "重要度": 4, "数量": 12 }],
|
||||
"event_types": [{ "事件类型": "地缘政治", "数量": 8 }],
|
||||
"source_dist": [{ "来源": "investinglive.com", "文章数": 45 }]
|
||||
"pipeline": { "raw_total": 442, "processed": 442, "deduped": 111, "embedded": 112, "qdrant": 2952 },
|
||||
"importance": [
|
||||
{ "importance": 2, "count": 2 }, { "importance": 3, "count": 31 },
|
||||
{ "importance": 4, "count": 26 }, { "importance": 5, "count": 5 }
|
||||
],
|
||||
"event_types": [
|
||||
{ "event_type": "宏观经济", "count": 15 }, { "event_type": "地缘政治", "count": 12 },
|
||||
{ "event_type": "财报披露", "count": 9 }, { "event_type": "央行决议", "count": 7 }
|
||||
],
|
||||
"source_dist": [
|
||||
{ "source": "InvestingLive", "count": 30 }, { "source": "MarketWatch", "count": 3 },
|
||||
{ "source": "CNBC", "count": 3 }, { "source": "Yahoo Finance", "count": 1 }
|
||||
]
|
||||
},
|
||||
"created_at": "2026-08-03T07:01:00"
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
> `stats` 为 JSON 快照:finance 含 `pipeline/sources/news/cninfo/xwlb`,intl 含 `pipeline/sentiment/importance/event_types/source_dist`。前端按 key 防御性读取(见 db_schema.md §3)。
|
||||
> `stats` 为 JSON 快照(示例为 2026-08-06 真实数据):finance 含 `pipeline/news/cninfo/xwlb`,intl 含 `pipeline/importance/event_types/source_dist`(另有 `sentiment`)。finance 与 intl 的 pipeline 内部 key 集**不同**(finance: `raw_total/raw_by_source/proc/dups/...`;intl: `raw_total/processed/deduped/embedded/qdrant`),前端按 key 防御性读取;各数字口径(`raw_total` ≠ `news.total` 等)见 `db_schema_v1.1.md` §3.1。
|
||||
|
||||
**传 `id`** → 单份详情(对象),增加 `events` 数组(按 `section, rank` 排序)。真实返回(id=182,58 条事件):
|
||||
|
||||
@@ -179,12 +198,15 @@ curl 'https://api.doorcome.cn/api/news/reports/?id=182'
|
||||
"title": "特朗普称已取消对伊朗的袭击计划,因双方就协议框架达成一致",
|
||||
"summary": "……",
|
||||
"sentiment": "neutral",
|
||||
"source": "investinglive.com",
|
||||
"source": "InvestingLive",
|
||||
"sources": ["InvestingLive"],
|
||||
"url": "https://www.investing.com/..."
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
> `sources` 为该新闻的**全部来源**(JSON 数组,字符串列表);`source` 为主来源(单值)。多源事件(如同一新闻被多家媒体转载)时 `sources` 含多个元素;历史事件可能为 `null`。
|
||||
|
||||
### curl 示例
|
||||
|
||||
```bash
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# CLAUDE-reference.md — 因子 + 表结构速查
|
||||
# 因子 + 表结构速查
|
||||
|
||||
## 因子速查(34 个,12 分类)
|
||||
|
||||
-1512
File diff suppressed because it is too large
Load Diff
+107
-681
File diff suppressed because it is too large
Load Diff
-675
@@ -1,675 +0,0 @@
|
||||
我认真看了你的目标和现有环境,我认为有一个关键点需要调整:
|
||||
|
||||
**不要把 Claude Code Plugin 当成系统主体。**
|
||||
|
||||
对于你的项目:
|
||||
|
||||
```text
|
||||
Claude Code
|
||||
MCP-Hub
|
||||
Django API
|
||||
AkShare
|
||||
MariaDB
|
||||
VectorBT
|
||||
Optuna
|
||||
LightGBM
|
||||
Qwen
|
||||
```
|
||||
|
||||
Claude Code 应该只是:
|
||||
|
||||
```text
|
||||
AI开发助手
|
||||
AI研究助手
|
||||
```
|
||||
|
||||
而不是:
|
||||
|
||||
```text
|
||||
系统运行时核心
|
||||
```
|
||||
|
||||
真正的核心应该是:
|
||||
|
||||
```text
|
||||
finance/
|
||||
```
|
||||
finance 目录下python虚拟环境位于 finance/.venv/,在项目根目录下 可用 source finance/.venv/bin/activate激活
|
||||
这个目录未来即使你不用 Claude、换成 Cursor、Codex、OpenHands、Aider,都应该能独立运行。
|
||||
|
||||
---
|
||||
|
||||
# 推荐总体架构
|
||||
|
||||
未来你的根目录:
|
||||
|
||||
```text
|
||||
/Users/summer/Downloads/cc-cursor
|
||||
|
||||
├── djapi/
|
||||
│
|
||||
├── finance/
|
||||
│
|
||||
├── mcp-servers/
|
||||
│
|
||||
├── shared/
|
||||
│
|
||||
├── docs/
|
||||
│
|
||||
└── .claude/
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 各目录职责
|
||||
|
||||
## djapi
|
||||
|
||||
仅负责:
|
||||
|
||||
```text
|
||||
数据库
|
||||
用户管理
|
||||
任务管理
|
||||
API接口
|
||||
报告管理
|
||||
```
|
||||
|
||||
类似:
|
||||
|
||||
```text
|
||||
Quant Platform Backend
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## finance
|
||||
|
||||
核心量化引擎
|
||||
|
||||
未来90%的代码都在这里。
|
||||
|
||||
---
|
||||
|
||||
结构:
|
||||
|
||||
```text
|
||||
finance/
|
||||
|
||||
├── config/
|
||||
│
|
||||
├── data/
|
||||
│
|
||||
├── factors/
|
||||
│
|
||||
├── models/
|
||||
│
|
||||
├── strategy/
|
||||
│
|
||||
├── optimizer/
|
||||
│
|
||||
├── backtest/
|
||||
│
|
||||
├── portfolio/
|
||||
│
|
||||
├── execution/
|
||||
│
|
||||
├── reports/
|
||||
│
|
||||
├── scheduler/
|
||||
│
|
||||
└── cli/
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 第一阶段(V1)
|
||||
|
||||
目标:
|
||||
|
||||
```text
|
||||
数据获取
|
||||
因子计算
|
||||
回测
|
||||
参数优化
|
||||
```
|
||||
|
||||
技术栈:
|
||||
|
||||
```text
|
||||
AkShare
|
||||
MariaDB
|
||||
VectorBT
|
||||
Optuna
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 实施周期
|
||||
|
||||
### Week 1
|
||||
|
||||
基础设施
|
||||
|
||||
---
|
||||
|
||||
目录:
|
||||
|
||||
```text
|
||||
finance/
|
||||
|
||||
config/
|
||||
data/
|
||||
database/
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
完成:
|
||||
|
||||
### DataManager
|
||||
|
||||
```python
|
||||
class DataManager:
|
||||
```
|
||||
|
||||
统一管理:
|
||||
|
||||
```text
|
||||
股票列表
|
||||
日线
|
||||
分钟线
|
||||
财务数据
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
不要让策略直接调用:
|
||||
|
||||
```python
|
||||
ak.stock_zh_a_hist()
|
||||
```
|
||||
|
||||
而是:
|
||||
|
||||
```python
|
||||
data_manager.get_daily()
|
||||
```
|
||||
|
||||
这样以后换 TuShare 不改策略。
|
||||
|
||||
---
|
||||
|
||||
### Week 2
|
||||
|
||||
因子引擎
|
||||
|
||||
建立:
|
||||
|
||||
```text
|
||||
factors/
|
||||
|
||||
technical/
|
||||
fundamental/
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
例如:
|
||||
|
||||
```text
|
||||
MomentumFactor
|
||||
RSIFactor
|
||||
ROEFactor
|
||||
PEFactor
|
||||
```
|
||||
|
||||
统一接口:
|
||||
|
||||
```python
|
||||
factor.calculate(df)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Week 3
|
||||
|
||||
VectorBT回测层
|
||||
|
||||
建立:
|
||||
|
||||
```text
|
||||
backtest/vectorbt/
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
统一接口:
|
||||
|
||||
```python
|
||||
engine.run(strategy)
|
||||
```
|
||||
|
||||
以后:
|
||||
|
||||
```text
|
||||
VectorBT
|
||||
Backtrader
|
||||
Zipline
|
||||
```
|
||||
|
||||
都能替换。
|
||||
|
||||
---
|
||||
|
||||
### Week 4
|
||||
|
||||
Optuna优化层
|
||||
|
||||
建立:
|
||||
|
||||
```text
|
||||
optimizer/
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
统一接口:
|
||||
|
||||
```python
|
||||
optimizer.optimize(strategy)
|
||||
```
|
||||
|
||||
实现:
|
||||
|
||||
```text
|
||||
MA
|
||||
RSI
|
||||
MACD
|
||||
```
|
||||
|
||||
自动寻优。
|
||||
|
||||
---
|
||||
|
||||
# 第二阶段(V2)
|
||||
|
||||
目标:
|
||||
|
||||
```text
|
||||
机器学习选股
|
||||
```
|
||||
|
||||
技术:
|
||||
|
||||
```text
|
||||
LightGBM
|
||||
CatBoost
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 新增目录
|
||||
|
||||
```text
|
||||
models/
|
||||
|
||||
├── lightgbm/
|
||||
└── catboost/
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
建立统一模型接口:
|
||||
|
||||
```python
|
||||
class BaseModel:
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
例如:
|
||||
|
||||
```python
|
||||
fit()
|
||||
|
||||
predict()
|
||||
|
||||
save()
|
||||
|
||||
load()
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
所有模型遵循:
|
||||
|
||||
```python
|
||||
BaseModel
|
||||
```
|
||||
|
||||
接口。
|
||||
|
||||
---
|
||||
|
||||
# 第三阶段(V3)
|
||||
|
||||
目标:
|
||||
|
||||
```text
|
||||
新闻因子
|
||||
公告因子
|
||||
研报因子
|
||||
```
|
||||
|
||||
技术:
|
||||
|
||||
```text
|
||||
Qwen
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
新增:
|
||||
|
||||
```text
|
||||
factors/sentiment/
|
||||
|
||||
models/qwen/
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
Qwen职责:
|
||||
|
||||
```text
|
||||
文本转因子
|
||||
```
|
||||
|
||||
例如:
|
||||
|
||||
```text
|
||||
公告
|
||||
↓
|
||||
Qwen
|
||||
↓
|
||||
sentiment_score
|
||||
↓
|
||||
feature_101
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
注意:
|
||||
|
||||
Qwen不是策略。
|
||||
|
||||
Qwen是:
|
||||
|
||||
```text
|
||||
因子生产工具
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 第四阶段(V4)
|
||||
|
||||
Agent化
|
||||
|
||||
新增:
|
||||
|
||||
```text
|
||||
finance/agents/
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
结构:
|
||||
|
||||
```text
|
||||
research_agent
|
||||
selection_agent
|
||||
risk_agent
|
||||
report_agent
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
例如:
|
||||
|
||||
### ResearchAgent
|
||||
|
||||
负责:
|
||||
|
||||
```text
|
||||
发现新因子
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### SelectionAgent
|
||||
|
||||
负责:
|
||||
|
||||
```text
|
||||
股票打分
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### RiskAgent
|
||||
|
||||
负责:
|
||||
|
||||
```text
|
||||
仓位控制
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Claude Code 接入时间点
|
||||
|
||||
很多人一开始就写 Skill。
|
||||
|
||||
我建议:
|
||||
|
||||
### 不要现在写大量 Skill
|
||||
|
||||
第一阶段只保留:
|
||||
|
||||
```text
|
||||
.claude/
|
||||
|
||||
skills/
|
||||
|
||||
factor-research
|
||||
backtest
|
||||
stock-selection
|
||||
```
|
||||
|
||||
三个就够。
|
||||
|
||||
---
|
||||
|
||||
等 V2 完成以后再扩展。
|
||||
|
||||
---
|
||||
|
||||
# 数据流设计(必须遵守)
|
||||
|
||||
未来所有代码都遵守:
|
||||
|
||||
```text
|
||||
Data
|
||||
↓
|
||||
|
||||
Factor
|
||||
↓
|
||||
|
||||
Model
|
||||
↓
|
||||
|
||||
Strategy
|
||||
↓
|
||||
|
||||
Backtest
|
||||
↓
|
||||
|
||||
Report
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
禁止:
|
||||
|
||||
```text
|
||||
Strategy
|
||||
↓
|
||||
直接访问AkShare
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
禁止:
|
||||
|
||||
```text
|
||||
Model
|
||||
↓
|
||||
直接访问数据库
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
全部通过 Service 层。
|
||||
|
||||
---
|
||||
|
||||
# 推荐实施顺序
|
||||
|
||||
## Sprint 1(1~2周)
|
||||
|
||||
完成:
|
||||
|
||||
```text
|
||||
finance/
|
||||
DataManager
|
||||
MariaDB
|
||||
AkShare
|
||||
```
|
||||
|
||||
目标:
|
||||
|
||||
```text
|
||||
获取任意股票历史数据
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Sprint 2(1周)
|
||||
|
||||
完成:
|
||||
|
||||
```text
|
||||
Factor Engine
|
||||
```
|
||||
|
||||
目标:
|
||||
|
||||
```text
|
||||
计算10个基础因子
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Sprint 3(1周)
|
||||
|
||||
完成:
|
||||
|
||||
```text
|
||||
VectorBT
|
||||
```
|
||||
|
||||
目标:
|
||||
|
||||
```text
|
||||
回测单因子策略
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Sprint 4(1周)
|
||||
|
||||
完成:
|
||||
|
||||
```text
|
||||
Optuna
|
||||
```
|
||||
|
||||
目标:
|
||||
|
||||
```text
|
||||
自动参数优化
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Sprint 5(2周)
|
||||
|
||||
完成:
|
||||
|
||||
```text
|
||||
LightGBM
|
||||
CatBoost
|
||||
```
|
||||
|
||||
目标:
|
||||
|
||||
```text
|
||||
训练未来5日收益预测模型
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Sprint 6(2周)
|
||||
|
||||
完成:
|
||||
|
||||
```text
|
||||
Qwen情绪因子
|
||||
```
|
||||
|
||||
目标:
|
||||
|
||||
```text
|
||||
新闻情绪打分
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Sprint 7
|
||||
|
||||
完成:
|
||||
|
||||
```text
|
||||
Agent
|
||||
Dashboard
|
||||
自动日报
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
按照这个路线,你的代码会从:
|
||||
|
||||
```text
|
||||
Mac Mini 单机量化研究平台
|
||||
```
|
||||
|
||||
平滑演进到:
|
||||
|
||||
```text
|
||||
多因子 + ML + LLM + Agent
|
||||
量化研究平台
|
||||
```
|
||||
|
||||
中间不会出现“推倒重写”的情况。最关键的是先把 **DataManager → Factor Engine → Backtest Engine → Model Engine** 四个基础引擎设计好,后面的 LightGBM、Qwen、Agent 都只是插件式增加能力。
|
||||
|
||||
@@ -1,5 +0,0 @@
|
||||
# 项目级 Reasonix 配置覆盖(仅此工作区生效)
|
||||
# 本机未安装 bubblewrap,为执行运维命令关闭 bash 沙箱(全局配置仍为 enforce)
|
||||
[sandbox]
|
||||
bash = "off"
|
||||
|
||||
Reference in New Issue
Block a user