diff --git a/.claude/skills/batch-sync.md b/.claude/skills/batch-sync.md deleted file mode 100644 index 36cb60d..0000000 --- a/.claude/skills/batch-sync.md +++ /dev/null @@ -1,26 +0,0 @@ -# batch-sync skill - -批量预热股票数据到 DB 缓存。 - -## 触发 - -用户说:预热缓存 / 同步数据 / warmup / batch sync / 补齐数据 / 全量同步 - -## 执行 - -```bash -cd finance && python cli/agent_cli.py warmup 50 -``` - -## 说明 - -- 每批 50 只股票,依次执行 `dm.sync_daily()` -- 已缓存 + 最新的 → 0 条跳过(增量) -- 未缓存 → Tushare 优先 → AkShare fallback -- 范围由 `.env` 中 `SENTIMENT_SCOPE_TYPE` + `SENTIMENT_SCOPE_INDEXES` 决定(默认沪深300+中证500) -- 可多次执行直到覆盖率 100% - -## 参数 - -`python cli/agent_cli.py warmup [N]` -- N: 每批股票数,默认 50。网络稳定时可调大到 100 diff --git a/README.md b/README.md index cc1e573..7b4a95a 100644 --- a/README.md +++ b/README.md @@ -7,225 +7,123 @@ ``` cc-cursor/ ├── finance/ # 核心量化引擎 -│ ├── config/ # 全局配置(MariaDB / AkShare) -│ ├── database/ # ORM 模型 + DAO(mac_ 前缀表) +│ ├── config/ # 全局配置 +│ ├── database/ # ORM 模型 + DAO │ ├── data/ # DataManager 统一数据层 -│ ├── factors/ # 因子引擎(34 因子 / 12 分类,含情绪因子) -│ ├── backtest/ # 回测引擎(VectorBT + 5 策略 + 截面回测) -│ ├── optimizer/ # Optuna 参数优化 + Walk-Forward +│ ├── factors/ # 因子引擎(34 因子 / 12 分类) +│ ├── backtest/ # 回测引擎(VectorBT + 5 策略) +│ ├── optimizer/ # Optuna 参数优化 │ ├── models/ # LightGBM / CatBoost ML 模型 -│ ├── agents/ # Agent 系统(4 Agent + 编排器 + CLI) +│ ├── agents/ # Agent 系统(4 Agent + 编排器) │ ├── cli/ # 命令行 & 验证脚本 -│ ├── reports/ # 自动日报输出目录 -│ └── .env # 环境变量配置(API Key / 分析范围) -├── djapi/ # Django API 后端(A 股数据 + 新闻联播 + 日报查询) -├── mcp-servers/ # MCP Server(Serena,本机工具,git 忽略) -├── shared/ # 共享工具(SSH 隧道脚本) -├── docs/ # 文档 & 使用指南 -└── .claude/ # Claude Code 配置 +│ └── reports/ # 日报输出目录 +├── djapi/ # Django API 后端 +├── shared/script/ # SSH 隧道脚本 +└── docs/ # 项目文档 ``` ## 数据流 ``` -Agent 编排层 - ├── ResearchAgent ── 因子发现(IC/IC_IR 评估) - ├── SelectionAgent ─ 多因子打分 + ML 预测 - ├── RiskAgent ────── 仓位控制 + 风险预警 - └── ReportAgent ──── 自动日报生成 - -基础引擎层 - DataManager ──→ FactorEngine ──→ BaseStrategy ──→ VectorBTEngine ──→ BacktestReport - │ │ │ - │ FeatureEngine OptunaEngine - │ │ │ - └──────→ LightGBM/CatBoost ←────────┘ - -情绪增强层 - NewsSource(AkShare/DB/MCP) ──→ QwenClient ──→ SentimentFactor ──→ FactorEngine +Data → Factor → Model → Strategy → Backtest → Report ``` -全部通过 Service 层中转:策略不直连 AkShare,模型不直连数据库,Agent 不重建引擎。 - ---- - -## 开发进度 - -| Sprint | 模块 | 关键成果 | 状态 | -|--------|------|----------|------| -| Sprint 0 | 基础设施 | DataManager + MariaDB 3 表 | ✅ | -| Sprint 1 | 因子引擎 | 34 因子 / 12 分类 | ✅ | -| Sprint 2 | 回测引擎 | VectorBT + 5 策略 + 截面回测 | ✅ | -| Sprint 3 | 参数优化 | Optuna + Walk-Forward | ✅ | -| Sprint 4 | ML 模型 | LightGBM + CatBoost + 特征工程 | ✅ | -| Sprint 5 | 情绪因子 | Qwen + 三源新闻聚合 + 日期对齐 | ✅ | -| Sprint 6 | Agent 系统 | 4 Agent + 编排器 + CLI + 自动日报 | ✅ | -| Sprint 7 | djapi API | 日报查询 ×2(news/reports + news/events) | ✅ | - -**全部 8 个 Sprint 已完成。** - ---- - -## 功能模块 - -### 数据层 `finance/data/` - -```python -from data.data_manager import DataManager -dm = DataManager(); dm.init_db() -stocks = dm.get_stock_list() # → 5,524 只 -daily = dm.get_daily("000001.SZ") # → 日线 -fina = dm.get_financial("000001.SZ") # → 财务数据 -dm.sync_daily("000001.SZ") # → 增量同步 -``` - -### 因子引擎 `finance/factors/` - -```python -from factors.registry import get_factor, list_factors -from factors.engine import FactorEngine - -engine = FactorEngine(dm) -factors = [get_factor("momentum_20"), get_factor("rsi_14")] -factor_df = engine.compute("000001.SZ", factors) -# → 34 个注册因子,12 个分类(动量/RSI/MACD/量价/布林/ATR/均线/波动率/换手率/振幅/基本面/情绪) -``` - -### 回测引擎 `finance/backtest/` - -```python -from backtest.vectorbt.engine import VectorBTEngine -from backtest.strategies.rsi_mean_revert import RSIMeanRevertStrategy - -engine_bt = VectorBTEngine(initial_capital=100_000, commission=0.0003) -report = engine_bt.run(RSIMeanRevertStrategy(oversold=30, overbought=70), price_df, factor_df) -# → 收益=29.4% 年化=4.3% 回撤=-19.1% 夏普=0.37 胜率=77.1% -``` - -5 个内置策略 + 自定义策略接口 + 截面回测 + BacktestReport 标准化报告。 - -### 参数优化 `finance/optimizer/` - -```python -from optimizer.engine import OptunaEngine -from optimizer.space import rsi_revert_space - -result = OptunaEngine(engine_bt).optimize( - RSIMeanRevertStrategy, rsi_revert_space, price_df, factor_df, - metric="sharpe", n_trials=200, -) -# → 最优参数: oversold=13, overbought=66 -# → 夏普: 0.37→0.60 (+62%), 回撤: -19.1%→-1.8% (10倍改善) -``` - -7 种优化目标 + 4 个预置搜索空间 + Walk-Forward 滚动验证 + 快捷函数。 - -### ML 模型 `finance/models/` - -```python -from models.features import FeatureEngine -from models.lightgbm.model import LightGBMModel - -fe = FeatureEngine(lookahead=5) -X, y = fe.build(factor_df, price_df, fit=True) -model = LightGBMModel(params={"n_estimators": 200}).fit(X_train, y_train) -pred = model.predict(X_test) # → IC 评估 + 特征重要性 + 交叉验证 + ML 策略回测 -``` - -Winsorize → 缺失填充 → RobustScaler → LightGBM/CatBoost 训练 → MLBenchmark 对比。 - -### 情绪因子 `finance/factors/sentiment/` - -```python -from factors.sentiment.sentiment_engine import SentimentEngine - -sent = SentimentEngine(dm) -sent_df = sent.compute("000001.SZ", max_news=20) -# → news_sent_5, news_conf_5, sent_delta_5 -``` - -三数据源聚合(AkShare 个股新闻 + 新闻联播 DB + MCP trendradar-news)、日期对齐(非交易日→最近交易日)、xwlb 偏移(昨日新闻→今日使用)、DashScope + Ollama 双后端。 - -### Agent 系统 `finance/agents/` - -```bash -python finance/cli/agent_cli.py daily # 完整每日流程 -python finance/cli/agent_cli.py picks 15 # 选股 Top 15 -python finance/cli/agent_cli.py risk # 风险评估 -python finance/cli/agent_cli.py research # 因子研究 -python finance/cli/agent_cli.py report # 生成日报 -``` - -4 个 Agent(Research/Selection/Risk/Report)+ 编排器 + 自动日报(reports/daily_YYYYMMDD.md)。 - ---- - ## 快速开始 ```bash -# SSH 隧道 +# 环境 & SSH 隧道 +conda activate quant bash shared/script/autossh.sh -# Python 环境 -conda activate quant # Python 3.11.13 - # 每日 Agent 运行 python finance/cli/agent_cli.py daily ``` -### 验证脚本 +## CLI 命令 ```bash -python finance/cli/demo_data_manager.py # Sprint 0 — DataManager -python finance/cli/demo_factor_engine.py # Sprint 1 — 因子引擎 -python finance/cli/demo_backtest.py # Sprint 2 — 回测引擎 -python finance/cli/demo_optimizer.py # Sprint 3 — 参数优化 -python finance/cli/demo_ml.py # Sprint 4 — ML 模型 -python finance/cli/demo_sentiment.py # Sprint 5 — 情绪因子 -python finance/cli/demo_sentiment_detail.py # Sprint 5 — 情绪因子(单股详情) +python finance/cli/agent_cli.py daily # 5 步完整流程 +python finance/cli/agent_cli.py picks 15 # 选股 Top 15 +python finance/cli/agent_cli.py risk # 风险评估 +python finance/cli/agent_cli.py research # 因子研究 +python finance/cli/agent_cli.py report 20260603 # 生成日报 +python finance/cli/agent_cli.py warmup 50 # 首次预热缓存 ``` ---- +## 验证脚本 + +```bash +python finance/cli/demo_data_manager.py --ts_code 600519.SH +python finance/cli/demo_factor_engine.py --ts_code 300750.SZ +python finance/cli/demo_backtest.py --ts_code 000001.SZ +python finance/cli/demo_optimizer.py --ts_code 000001.SZ --trials 100 +python finance/cli/demo_ml.py --ts_code 000001.SZ --lookahead 5 +python finance/cli/demo_sentiment.py --ts_code 600519.SH +python finance/cli/demo_sentiment_detail.py --ts_code 600519.SH --date 20260603 +``` ## 技术栈 -| 组件 | 技术 | 版本 | 状态 | -|------|------|------|------| -| 数据获取 | AkShare | 1.18.64 | ✅ | -| 数据库 | MariaDB (SSH 隧道) | — | ✅ | -| 因子/特征 | pandas / numpy / sklearn | 2.3 / 2.0 / 1.9 | ✅ | -| 回测引擎 | VectorBT | 1.0 | ✅ | -| 参数优化 | Optuna | 4.9 | ✅ | -| ML 模型 | LightGBM / CatBoost | 4.6 / 1.2 | ✅ | -| NLP 情绪 | Qwen (DashScope / Ollama) | turbo / 2.5 | ✅ | -| Agent 框架 | 自研编排器 | — | ✅ | -| API 后端 | Django + uWSGI | 5.2 | 已有 | -| 代码分析 | Serena MCP | — | 本机工具(不随仓库分发) | +| 组件 | 技术 | +|------|------| +| 数据获取 | AkShare + Tushare (双源) | +| 数据库 | MariaDB (SSH 隧道) | +| 因子/特征 | pandas / numpy / sklearn | +| 回测引擎 | VectorBT 1.0 | +| 参数优化 | Optuna 4.9 | +| ML 模型 | LightGBM 4.6 + CatBoost 1.2 | +| NLP 情绪 | Qwen (DashScope / Ollama) | +| Agent 编排 | 自研编排器 | +| API 后端 | Django 5.2 + uWSGI | ---- +## 开发进度 -## 设计原则 +| Sprint | 模块 | 状态 | +|--------|------|------| +| Sprint 0 | 基础设施(DataManager + MariaDB) | ✅ | +| Sprint 1 | 因子引擎(34 因子 / 12 分类) | ✅ | +| Sprint 2 | VectorBT 回测(5 策略 + 截面) | ✅ | +| Sprint 3 | Optuna 优化(+ Walk-Forward) | ✅ | +| Sprint 4 | ML 模型(LightGBM + CatBoost) | ✅ | +| Sprint 5 | Qwen 情绪因子(三源新闻) | ✅ | +| Sprint 6 | Agent 系统(4 Agent + CLI) | ✅ | +| Sprint 7 | djapi API(日报查询 ×2) | ✅ | -- **模块隔离**:各引擎通过统一接口交互,可替换实现(VectorBT → Backtrader) -- **接口标准化**:因子 `calculate(df)→Series` / 策略 `generate_signals(df)→Series` / 模型 `fit/predict/save/load` / 优化 `optimize()→Result` -- **数据层统一**:策略/模型不直连数据源,全部通过 DataManager -- **Agent 不重建轮子**:Agent 通过依赖注入复用已有引擎,编排而非重建 -- **防前视偏差**:时间序列交叉验证、expanding window 统计量 -- **渐进演进**:全链路 8 个 Sprint 平滑推进,无推倒重写 +**全部 8 个 Sprint 已完成。** ## 文档 -- [使用指南](./docs/usage.md) — 详细使用说明(12 章节,含代码示例) -- [使用指南 (HTML)](./docs/usage.html) — 网页版使用指南 -- [新闻日报 API](./docs/news_report_api.md) — djapi 日报查询接口使用手册 +| 文档 | 内容 | +|------|------| +| [使用指南](docs/usage.md) | 各模块使用方法和代码示例 | +| [架构说明](docs/architecture.md) | 项目架构、数据流、设计原则 | +| [开发指南](docs/development.md) | 环境搭建、开发约定、模块说明 | +| [部署说明](docs/deployment.md) | 本地环境、服务器、uWSGI、rsync 部署 | +| [因子与表结构速查](docs/reference.md) | 34 因子注册表、DB 表结构 | +| [数据层详解](docs/data-layer.md) | DataManager、数据库、缓存策略 | +| [因子引擎详解](docs/factors.md) | 因子计算、情绪引擎、新闻源 | +| [回测引擎详解](docs/backtest.md) | VectorBT、策略、信号工具、Optuna | +| [ML 模型详解](docs/ml-models.md) | 特征工程、LightGBM/CatBoost | +| [Agent 系统详解](docs/agents.md) | Agent 架构、CLI、日报 | +| [DJAPI 接口](docs/api.md) | Django API 端点参考 | +| [日报查询 API](docs/news_report_api.md) | news/reports + news/events 接口 | +| [日报数据库](docs/db_schema_v1.1.md) | news_report / news_event 表结构 | ## 子项目 -- [djapi](./djapi/README.md) — Django API 后端:A 股数据 API(16 端点)+ 新闻联播处理 + 日报查询(news/reports、news/events) +- [djapi](djapi/README.md) — Django API 后端:A 股数据 API(16 端点)+ 新闻联播处理 + 日报查询 + +## 设计原则 + +- **模块隔离**:各引擎通过统一接口交互,可替换实现 +- **接口标准化**:因子 `calculate(df)→Series` / 策略 `generate_signals(df)→Series` / 模型 `fit/predict/save/load` +- **数据层统一**:策略/模型不直连数据源,全部通过 DataManager +- **Agent 不重建轮子**:Agent 通过依赖注入复用已有引擎 +- **防前视偏差**:时间序列交叉验证、expanding window 统计量 ## 数据库连接 ```bash bash shared/script/autossh.sh # host: 127.0.0.1:13306 user: myquant database: myquant table_prefix: mac_ -``` +``` \ No newline at end of file diff --git a/continuation.md b/continuation.md deleted file mode 100644 index 1d76a09..0000000 --- a/continuation.md +++ /dev/null @@ -1,83 +0,0 @@ -# continuation.md — cc-cursor 项目状态 - -生成时间:2026-06-07(全部 Sprint 完成 + 生产加固 + djapi 数据源归一化 + Git 初始化) - ---- - -## Git 状态 - -- 仓库:https://github.com/Simon2046/myquant -- 分支:`main` -- commit:`271a934` — Initial commit: cc-cursor 全链路量化研究平台 -- 文件:293 个文件,59,598 行 -- 已排除:`.env`、`mcp-servers/serena`、`__pycache__`、`.parquet`、`.db` - ---- - -## 全部 Sprint 完成 ✅ - -| Sprint | 模块 | 状态 | -|--------|------|------| -| 0 | 基础设施(DataManager + MariaDB) | ✅ | -| 1 | 因子引擎(34 因子 / 12 分类) | ✅ | -| 2 | VectorBT 回测(5 策略 + 截面 + BacktestReport) | ✅ | -| 3 | Optuna 优化(+ Walk-Forward) | ✅ | -| 4 | ML 模型(LightGBM + CatBoost + MLStrategy) | ✅ | -| 5 | Qwen 情绪因子(三源新闻 + 日期对齐) | ✅ | -| 6 | Agent 系统(4 Agent + CLI + 日报 .md/.html) | ✅ | - ---- - -## 生产稳定性加固(14 项) - -| # | 项 | 文件 | -|---|-----|------| -| 1 | Tushare 双数据源(优先) | `finance/data/data_manager.py` | -| 2 | 指数 vs 个股自动路由 | `finance/data/sources/akshare_source.py` | -| 3 | SSH 自动恢复(多次重连 + pool_pre_ping) | `finance/database/connection.py` | -| 4 | save_daily 先删后插(防主键冲突) | `finance/database/dao.py` | -| 5 | load_dotenv 绝对路径 + 模块加固 | `finance/config/settings.py` + 3 文件 | -| 6 | 日报 5d/20d 修复(idx=-1→pos=len-1) | `finance/agents/report_agent.py` | -| 7 | RiskAgent 改用上证指数 | `finance/agents/risk_agent.py` | -| 8 | 日报增加"昨日对比" + 数据截止 | `finance/agents/report_agent.py` | -| 9 | mac_report 表 utf8mb4 + DATE + DATETIME | `finance/database/models.py` | -| 10 | 日报自动存入 DB + emoji 兼容 | `finance/reports/storage.py` + 8 CLI | -| 11 | CLAUDE-*.md 文档化 9 条已知 Bug | `CLAUDE-data.md` + `CLAUDE-agents.md` | -| 12 | demo 脚本全参数化 | `finance/cli/demo_*.py` | -| 13 | **djapi 数据源归一化(10→1 入口)** | `djapi/api/stock/data_source.py` | -| 14 | **indexDatas API 参数修正 + 容错** | `djapi/api/views.py` + `getIndexs.py` | - ---- - -## djapi 数据源归一化 - -- 新增 `djapi/api/stock/data_source.py` — 统一入口 - - `get_tushare_pro()` — 全局单例(线程安全) - - `get_daily()` — 双源 fallback (Tushare→AkShare) - - `get_mysql_db()` — MySQL 全局单例 - - Token 兼容 `TUSHARE_TS_TOKEN` / `TUSHARE_TOKEN` -- 10 个模块迁移完成 -- `getDivData_AK.py` 标记废弃 -- `getIndexs.py` 修复:`index_dailybasic` 失败不阻塞,异常 raise 而非静默返回空 -- `views.py` 修正:`indexDatas` 参数 `index_name` → `tscode`,描述从"股票代码"→"指数代码",新增 `_PARAM_INDEX_CODE` -- 已部署到 `api.doorcome.cn` ✅ - ---- - -## CLI 命令 - -```bash -agent_cli.py daily / picks / risk / research / report / warmup -demo_*.py(全部支持 --ts_code --date 等参数) -``` - -## 文档 - -- `CLAUDE.md` — 入口 + 路由 + 多步任务规则 -- `CLAUDE-data.md` — 数据层 + 5 条已知 Bug -- `CLAUDE-factors.md` — 因子引擎 -- `CLAUDE-backtest.md` — 回测 + 优化 -- `CLAUDE-ml.md` — ML 模型 -- `CLAUDE-agents.md` — Agent + CLI + 4 条已知 Bug -- `CLAUDE-reference.md` — 因子/表结构速查 -- `docs/usage.md` + `docs/usage.html` — 使用指南 diff --git a/djapi/.env.example b/djapi/.env.example index 6c373a7..44b573a 100644 --- a/djapi/.env.example +++ b/djapi/.env.example @@ -13,7 +13,7 @@ MYSQL_PASSWORD=your-mysql-password MYSQL_DATABASE=myquant # 日报结构化入库 (news_report / news_event) 只读查询 -# 与 report_db_design.md §7 一致;密码必填,缺失时接口直接报错 +# 与 docs/db_schema_v1.1.md 一致;密码必填,缺失时接口直接报错 NEWS_DB_HOST=127.0.0.1 NEWS_DB_PORT=3306 NEWS_DB_USER=myquant diff --git a/djapi/.mcp.json b/djapi/.mcp.json deleted file mode 100644 index b652839..0000000 --- a/djapi/.mcp.json +++ /dev/null @@ -1,16 +0,0 @@ -{ - "mcpServers": { - "serena-djapi": { - "command": "uv", - "args": [ - "run", - "--directory", - "/Users/summer/Downloads/cc-cursor/mcp-servers/serena", - "serena", - "start-mcp-server", - "--project", - "/Users/summer/Downloads/cc-cursor/djapi" - ] - } - } -} diff --git a/djapi/.serena/.gitignore b/djapi/.serena/.gitignore deleted file mode 100644 index 2e510af..0000000 --- a/djapi/.serena/.gitignore +++ /dev/null @@ -1,2 +0,0 @@ -/cache -/project.local.yml diff --git a/djapi/.serena/memories/code_style_and_conventions.md b/djapi/.serena/memories/code_style_and_conventions.md deleted file mode 100644 index 4a363fe..0000000 --- a/djapi/.serena/memories/code_style_and_conventions.md +++ /dev/null @@ -1,23 +0,0 @@ -# Code Style & Conventions - -## Python -- Django app: all business logic in `api/stock/`, not in views -- views.py is thin forwarding layer: extract params -> call function -> return Response -- Double import pattern for standalone scripts: try relative import first, fall back to absolute -- Use `viewFunc_tsCodeAndDate()` wrapper for ts_code + date_range endpoints -- Use `viewFunc_singleParam()` wrapper for single-param endpoints -- DRF `@api_view(['GET'])` + `@extend_schema` on all views -- DRF `Response` (not `JsonResponse`) — no `safe=False` parameter -- Configuration split: config.py (token) / strategy_config.py / scan_config.py - -## Constraints -- `api/video/` is protected — do NOT modify unless user explicitly asks -- No python-dotenv dependency — use stdlib env loaders only -- Backward compatibility: keep re-exports when splitting modules -- Server `.env` file manages all secrets; uwsgi.ini only has DJANGO_SETTINGS_MODULE - -## Secrets -- All API keys/tokens/passwords via os.getenv() -- Local: .env file (not committed) -- Server: /home/simon/myquant/djapi/.env -- Django loads via djapi/env_loader.py, video loads via api/video/env.py diff --git a/djapi/.serena/memories/project_overview.md b/djapi/.serena/memories/project_overview.md deleted file mode 100644 index dd116ba..0000000 --- a/djapi/.serena/memories/project_overview.md +++ /dev/null @@ -1,32 +0,0 @@ -# Project Overview - -djapi is a Django 5.2 project providing financial data APIs for A-share stocks and CCTV news broadcast video processing. - -## Tech Stack -- Python 3.10, Django 5.2, uWSGI, nginx -- Tushare (stock data), akshare (alternative stock data) -- DRF + drf-spectacular (API documentation) -- MySQL (business data), SQLite (Django admin only) -- yt-dlp + ffmpeg + pydub (video/audio processing) -- DashScope (ASR), DeepSeek API (AI text processing) - -## Architecture -- Single Django app: `api` -- `api/stock/` — stock data module (Tushare/akshare -> pandas -> JsonResponse/DRF Response) -- `api/video/` — independent video processing pipeline (download -> audio -> ASR -> AI split -> MySQL) -- views.py is thin: extracts params, calls stock functions, returns Response - -## Key Files -- `api/views.py` — all ~15 API views, using @api_view + @extend_schema -- `api/stock/stock_utils.py` — shared utilities: tscodeCheck, viewFunc_tsCodeAndDate, viewFunc_singleParam -- `api/stock/config.py` — Tushare token + re-exports from strategy_config, scan_config -- `api/serializers.py` — 13 DRF Serializer classes -- `djapi/env_loader.py` — .env file loader (stdlib, no python-dotenv) -- `api/video/env.py` — standalone .env loader for video module -- `api/utils/mysql_handler.py` — shared MySQLDB class - -## Deployment -- Server: simon@doorcome.cn, path: /home/simon/myquant/djapi/ -- Virtual env: /opt/miniconda/envs/django/ -- uWSGI on port 5004, nginx reverse proxy -- Domains: api.doorcome.cn, echart.doorcome.cn diff --git a/djapi/.serena/memories/suggested_commands.md b/djapi/.serena/memories/suggested_commands.md deleted file mode 100644 index c5fabaf..0000000 --- a/djapi/.serena/memories/suggested_commands.md +++ /dev/null @@ -1,44 +0,0 @@ -# Suggested Commands - -## Development -```bash -python manage.py runserver 0.0.0.0:8000 # dev server -python manage.py check --deploy # check config -python manage.py test api # run tests -``` - -## uWSGI -```bash -uwsgi --ini uwsgi.ini # start -uwsgi --reload uwsgi.pid # hot reload -uwsgi --stop uwsgi.pid # stop -# On server: -/opt/miniconda/envs/django/bin/uwsgi --ini /home/simon/myquant/djapi/uwsgi.ini -kill $(lsof -ti:5004) # force stop -``` - -## Deploy -```bash -# Full sync (exclude production data) -rsync -avz --delete \ - --exclude='.env' --exclude='db.sqlite3' \ - --exclude='*.log' --exclude='uwsgi.pid' \ - --exclude='__pycache__/' --exclude='*.pyc' \ - --exclude='xwlb_video/' --exclude='audio_processing/' \ - /Users/summer/Downloads/cc-cursor/djapi/ \ - simon@doorcome.cn:/home/simon/myquant/djapi/ - -# Single file sync MUST use full target path -rsync -avz api/views.py simon@doorcome.cn:/home/simon/myquant/djapi/api/views.py -``` - -## API Docs -- /api/docs/ — Swagger UI -- /api/redoc/ — ReDoc -- /api/schema/ — OpenAPI JSON - -## Video Processing -```bash -cd api/video -python main.py -``` diff --git a/djapi/.serena/project.yml b/djapi/.serena/project.yml deleted file mode 100644 index 653c9f4..0000000 --- a/djapi/.serena/project.yml +++ /dev/null @@ -1,120 +0,0 @@ -# the name by which the project can be referenced within Serena -project_name: "djapi" - - -# list of languages for which language servers are started; choose from: -# al ansible bash clojure cpp -# cpp_ccls crystal csharp csharp_omnisharp dart -# elixir elm erlang fortran fsharp -# go groovy haskell haxe hlsl -# java json julia kotlin lean4 -# lua luau markdown matlab msl -# nix ocaml pascal perl php -# php_phpactor powershell python python_jedi python_ty -# r rego ruby ruby_solargraph rust -# scala solidity swift systemverilog terraform -# toml typescript typescript_vts vue yaml -# zig -# (This list may be outdated. For the current list, see values of Language enum here: -# https://github.com/oraios/serena/blob/main/src/solidlsp/ls_config.py -# For some languages, there are alternative language servers, e.g. csharp_omnisharp, ruby_solargraph.) -# Note: -# - For C, use cpp -# - For JavaScript, use typescript -# - For Free Pascal/Lazarus, use pascal -# Special requirements: -# Some languages require additional setup/installations. -# See here for details: https://oraios.github.io/serena/01-about/020_programming-languages.html#language-servers -# When using multiple languages, the first language server that supports a given file will be used for that file. -# The first language is the default language and the respective language server will be used as a fallback. -# Note that when using the JetBrains backend, language servers are not used and this list is correspondingly ignored. -languages: -- typescript -- python - -# the encoding used by text files in the project -# For a list of possible encodings, see https://docs.python.org/3.11/library/codecs.html#standard-encodings -encoding: "utf-8" - -# line ending convention to use when writing source files. -# Possible values: unset (use global setting), "lf", "crlf", or "native" (platform default) -# This does not affect Serena's own files (e.g. memories and configuration files), which always use native line endings. -line_ending: - -# The language backend to use for this project. -# If not set, the global setting from serena_config.yml is used. -# Valid values: LSP, JetBrains -# Note: the backend is fixed at startup. If a project with a different backend -# is activated post-init, an error will be returned. -language_backend: - -# whether to use project's .gitignore files to ignore files -ignore_all_files_in_gitignore: true - -# advanced configuration option allowing to configure language server-specific options. -# Maps the language key to the options. -# Have a look at the docstring of the constructors of the LS implementations within solidlsp (e.g., for C# or PHP) to see which options are available. -# No documentation on options means no options are available. -ls_specific_settings: {} - -# list of additional paths to ignore in this project. -# Same syntax as gitignore, so you can use * and **. -# Note: global ignored_paths from serena_config.yml are also applied additively. -ignored_paths: [] - -# whether the project is in read-only mode -# If set to true, all editing tools will be disabled and attempts to use them will result in an error -# Added on 2025-04-18 -read_only: false - -# list of tool names to exclude. -# This extends the existing exclusions (e.g. from the global configuration) -# Find the list of tools here: https://oraios.github.io/serena/01-about/035_tools.html -excluded_tools: [] - -# list of tools to include that would otherwise be disabled (particularly optional tools that are disabled by default). -# This extends the existing inclusions (e.g. from the global configuration). -# Find the list of tools here: https://oraios.github.io/serena/01-about/035_tools.html -included_optional_tools: [] - -# fixed set of tools to use as the base tool set (if non-empty), replacing Serena's default set of tools. -# This cannot be combined with non-empty excluded_tools or included_optional_tools. -# Find the list of tools here: https://oraios.github.io/serena/01-about/035_tools.html -fixed_tools: [] - -# list of mode names that are to be activated by default, overriding the setting in the global configuration. -# The full set of modes to be activated is base_modes (from global config) + default_modes + added_modes. -# If the setting is undefined/empty, the default_modes from the global configuration (serena_config.yml) apply. -# Otherwise, this overrides the setting from the global configuration (serena_config.yml). -# Therefore, you can set this to [] if you do not want the default modes defined in the global config to apply -# for this project. -# This setting can, in turn, be overridden by CLI parameters (--mode). -# See https://oraios.github.io/serena/02-usage/050_configuration.html#modes -default_modes: - -# list of mode names to be activated additionally for this project, e.g. ["query-projects"] -# The full set of modes to be activated is base_modes (from global config) + default_modes + added_modes. -# See https://oraios.github.io/serena/02-usage/050_configuration.html#modes -added_modes: - -# initial prompt for the project. It will always be given to the LLM upon activating the project -# (contrary to the memories, which are loaded on demand). -initial_prompt: "" - -# time budget (seconds) per tool call for the retrieval of additional symbol information -# such as docstrings or parameter information. -# This overrides the corresponding setting in the global configuration; see the documentation there. -# If null or missing, use the setting from the global configuration. -symbol_info_budget: - -# list of regex patterns which, when matched, mark a memory entry as read‑only. -# Extends the list from the global configuration, merging the two lists. -read_only_memory_patterns: [] - -# list of regex patterns for memories to completely ignore. -# Matching memories will not appear in list_memories or activate_project output -# and cannot be accessed via read_memory or write_memory. -# To access ignored memory files, use the read_file tool on the raw file path. -# Extends the list from the global configuration, merging the two lists. -# Example: ["_archive/.*", "_episodes/.*"] -ignored_memory_patterns: [] diff --git a/djapi/CLAUDE.md b/djapi/CLAUDE.md index e8136a8..1c8b0e9 100644 --- a/djapi/CLAUDE.md +++ b/djapi/CLAUDE.md @@ -99,7 +99,7 @@ python api/video/main.py ## 注意事项 -- `config.py` 中的 TS_TOKEN 和 `deepseek.py`/`ai.py`/`audioRead.py` 中的 API key、`mysqlHandle.py` 中的数据库密码均为硬编码 —— 生产环境应迁移到环境变量 +- 所有密钥已迁移到环境变量,通过 `.env` 统一管理 - `api/stock/` 下的模块支持两种导入方式(相对导入和绝对导入),这是为了兼容「作为 Django app 被调用」和「直接命令行运行脚本」两种场景 - `api/video/` 模块设计为独立命令行运行,不依赖 Django 框架 - `db.sqlite3` 已提交到代码库,包含 Django admin 的用户数据 diff --git a/djapi/README.md b/djapi/README.md index f224c6d..f5342cc 100644 --- a/djapi/README.md +++ b/djapi/README.md @@ -98,7 +98,7 @@ uwsgi --stop uwsgi.pid | `news/reports/` | report_type, start_date, end_date, id | 日报查询(默认最近 24h;传 id 返回单份详情含事件) | | `news/events/` | days, importance, report_type, section, limit | 重要事件聚合(跨日报,最近 N 天 importance≥阈值) | -日报查询接口详细说明见 [`docs/news_report_api.md`](../docs/news_report_api.md)(表结构见 `djapi/docs/db_schema.md`)。 +日报查询接口详细说明见 [`docs/news_report_api.md`](../docs/news_report_api.md)(表结构见 `../docs/db_schema_v1.1.md`)。 API 文档(Swagger):`/api/docs/` OpenAPI Schema:`/api/schema/` diff --git a/djapi/api/report/query.py b/djapi/api/report/query.py index a15f567..581536c 100644 --- a/djapi/api/report/query.py +++ b/djapi/api/report/query.py @@ -1,7 +1,7 @@ """ -news_report / news_event 只读查询层(日报结构化入库,见 docs/db_schema.md)。 +news_report / news_event 只读查询层(日报结构化入库,见 docs/db_schema_v1.1.md)。 -连接配置来自环境变量(与 docs/report_db_design.md §7 保持一致): +连接配置来自环境变量(与 docs/db_schema_v1.1.md 保持一致): NEWS_DB_HOST / NEWS_DB_PORT / NEWS_DB_USER / NEWS_DB_PASSWORD / NEWS_DB_NAME NEWS_DB_PASSWORD 缺失时直接报错,禁止默认密码。 @@ -45,6 +45,18 @@ def _connect(): return mysql.connector.connect(**load_db_config()) +def _parse_json(value): + """把 JSON 字符串列(如 sources)解析为 dict/list;已是对象则原样返回。""" + if value is None: + return None + if isinstance(value, str): + try: + return json.loads(value) + except (TypeError, ValueError): + return None + return value + + def _row_to_dict(row: dict) -> dict: """序列化行:stats JSON 解析、日期/时间转 ISO 字符串。""" d = dict(row) @@ -87,11 +99,14 @@ def fetch_reports( report = _row_to_dict(row) cur.execute( "SELECT id, section, rank, importance, event_type, title, " - "summary, sentiment, source, url " + "summary, sentiment, source, sources, url " "FROM news_event WHERE report_id = %s ORDER BY section, rank", (report_id,), ) - report["events"] = [dict(r) for r in cur.fetchall()] + report["events"] = [ + {**dict(r), "sources": _parse_json(r["sources"])} + for r in cur.fetchall() + ] return report where, params = [], [] @@ -153,7 +168,7 @@ def fetch_important_events( sql = ( "SELECT r.report_date, r.report_type, e.id, e.section, e.rank, " "e.importance, e.event_type, e.title, e.summary, e.sentiment, " - "e.source, e.url " + "e.source, e.sources, e.url " "FROM news_event e " "JOIN news_report r ON r.id = e.report_id " "WHERE " + " AND ".join(where) @@ -163,6 +178,10 @@ def fetch_important_events( ) params.append(int(limit)) cur.execute(sql, tuple(params)) - return [dict(r) for r in cur.fetchall()] + rows = cur.fetchall() + return [ + {**dict(r), "sources": _parse_json(r["sources"])} + for r in rows + ] finally: conn.close() diff --git a/djapi/api/report/serializers.py b/djapi/api/report/serializers.py index 6c9f4db..13162a1 100644 --- a/djapi/api/report/serializers.py +++ b/djapi/api/report/serializers.py @@ -14,6 +14,7 @@ class EventSerializer(serializers.Serializer): summary = serializers.CharField(allow_null=True) sentiment = serializers.CharField(allow_null=True) source = serializers.CharField(allow_null=True) + sources = serializers.JSONField(allow_null=True) url = serializers.CharField(allow_null=True) @@ -47,4 +48,5 @@ class ImportantEventSerializer(serializers.Serializer): summary = serializers.CharField(allow_null=True) sentiment = serializers.CharField(allow_null=True) source = serializers.CharField(allow_null=True) + sources = serializers.JSONField(allow_null=True) url = serializers.CharField(allow_null=True) diff --git a/djapi/api/report/tests.py b/djapi/api/report/tests.py index 74b8a9c..17c018e 100644 --- a/djapi/api/report/tests.py +++ b/djapi/api/report/tests.py @@ -122,6 +122,17 @@ class NewsEventsAPITest(TestCase): self.assertEqual(kwargs['report_type'], 'intl') self.assertEqual(kwargs['section'], 'intl') + @patch('api.report.query.fetch_important_events', + return_value=[{'id': 1, 'title': 'x', 'source': 'yicai', + 'sources': ['yicai', 'stcn']}]) + def test_sources_field_passthrough(self, mock_fetch): + """事件聚合响应原样透传 sources(JSON 数组)""" + resp = self.client.get(self.url) + self.assertEqual(resp.status_code, 200) + body = resp.json() + self.assertEqual(body[0]['sources'], ['yicai', 'stcn']) + self.assertEqual(body[0]['source'], 'yicai') + def test_invalid_days(self): resp = self.client.get(self.url, {'days': 'abc'}) self.assertEqual(resp.status_code, 400) diff --git a/djapi/continuation.md b/djapi/continuation.md deleted file mode 100644 index 08030f4..0000000 --- a/djapi/continuation.md +++ /dev/null @@ -1,135 +0,0 @@ -# continuation.md - -## 当前项目状态 - -djapi — Django 5.2 金融数据 API 项目,2026-06-17 已部署。 - -## Checkpoint 记录 - -| 日期 | 内容 | -|------|------| -| 2026-08-03 | 新增日报查询 API ×2(news/reports/ + news/events/),基于 news_report/news_event 表,已部署 doorcome ✅ | - -服务器:`simon@doorcome.cn`,路径 `/home/simon/myquant/djapi/`,虚拟环境 `/opt/miniconda/envs/django/`。 - -## 已完成 - -### 1. 安全:密钥统一管理 -- 所有密钥 → 环境变量,`.env` 统一管理 -- `djapi/env_loader.py`(Django 端)+ `api/video/env.py`(video 端)双加载器 -- 共享 `MySQLDB` → `api/utils/mysql_handler.py` -- `.env.example`、`.gitignore` - -### 2. 代码质量 -- `api/views.py`:227 → ~130 行,消除重复 -- `api/stock/stock_utils.py`:`viewFunc_singleParam()` 包装器 -- `api/stock/config.py`:拆分为 config / strategy_config / scan_config - -### 3. drf-spectacular 集成 -- 14 端点 `@api_view` + `@extend_schema`,8 tag 分组 -- 13 Serializer,Swagger `/api/docs/` - -### 4. 股息率 API 优化 -- **删除** `api/stock/getDivData_AK.py`(akshare 版),`/api/getdivak/` 路由移除 -- **优化** `api/stock/getStockDiv2.py`: - - TTM 计算:`calculate_ttm_div` 行级循环 O(n²) → `rolling('360D').sum()` O(n) - - 删除向前填充逻辑(~30 行),避免与毛刺平滑冲突 -- **修复** `api/stock/smoothBrush.py`:if/elif 分支中 prev_valid/next_valid 赋值反了 - -### 5. video 模块重构与 Bug 修复 -- 新增 `api/video/env.py` — .env 加载 -- `newsRedo.py` 重写 — 三分支智能重处理 -- **P0 修复**:`getVideo5.py` 日期校验 bug(`start_date > start_date` → `start_date > end_date`) -- **P1 清理**:删除 `ai.py`(两个函数均为死代码),清理 `newsProcess.py` 冗余 import -- **P2 修复**:`audioRead.py` — `transcribe_audio` 异常时返回 `['', '']` 统一类型;`analyze_and_correct_text` 防御 None -- **P3 修复**:`newsProcess.py` — `news_to_db()` JSON 解析自适应 dict/list(DeepSeek json_object 模式返回 dict 包装) - -### 6. 文档与测试 -- `CLAUDE.md`、`README.md`、`continuation.md` -- 18 个单元测试 - -## video 目录文件现状(10 个 .py) - -| 文件 | 职责 | -|------|------| -| `env.py` | .env 加载 | -| `getVideo5.py` | 主流程:抓取→下载→ASR→入库 | -| `audioRead.py` | 音频转换、分割、ASR 识别、文本纠错 | -| `deepseek.py` | DeepSeek API 封装(类 + `deepseek_text` 函数,支持 `response_format`) | -| `newsProcess.py` | AI 新闻分割+标题提取,JSON 自适应解析 | -| `newsRedo.py` | 手动重处理(三分支) | -| `main.py` | 定时任务入口(当天) | -| `main_videos.py` | 批量补缺(扫描缺失日期) | -| `mysqlHandle.py` | MySQLDB 重新导出 | - -## 所有 API 端点(16 个) - -| 端点 | 数据源 | 说明 | -|------|--------|------| -| `stockbasic/` | Tushare | 日线行情 | -| `stockinfo/` | Tushare | 个股基本信息 | -| `industrys/` | Tushare | 行业股票列表 | -| `stockparam/` | Tushare | 个股参数 | -| `stockep/` | Tushare | TTM EPS | -| `quarterlyEps/` | Tushare | 季度 EPS | -| `indexByName/` | Tushare | 指数查询 | -| `indexDatas/` | Tushare | 指数行情 | -| `dailymargin/` | Tushare | 每日融资融券汇总 | -| `stockmargin/` | Tushare | 个股融资融券 | -| `finance/` | Tushare | 财务报表分析 | -| `getdiv/` | Tushare | 股息率(TTM rolling + 毛刺平滑) | -| `xwlbNews/` | MySQL | 新闻联播原始文本 | -| `xwlbFine/` | MySQL | 新闻联播 AI 精编 | -| `news/reports/` | MySQL (news_) | 日报查询:默认最近 24h;传 id 返回详情含事件 | -| `news/events/` | MySQL (news_) | 重要事件聚合:最近 N 天 importance≥阈值 | - ---- - -## 日报查询 API(2026-08-03 新增) - -### 模块 - -- 新增 `api/report/` 包(独立于 stock):`query.py`(连库+查询 SQL)/ `views.py`(2 视图)/ `serializers.py`(OpenAPI)/ `tests.py`(17 个单测,mock 查询层) -- `api/urls.py` 注册 `news/reports/`、`news/events/`;`settings.py` SPECTACULAR TAGS 加「日报」 -- 数据库:doorcome 本机 MariaDB `myquant` 库 `news_report`(180 行)+ `news_event`(4372 条),与现有 `MYSQL_*` 同库同用户 -- 连接配置:服务器 `djapi/.env` 新增 `NEWS_DB_*`(复用 MYSQL_* 值,密码必填否则 500) -- 文档:`docs/news_report_api.md`(使用手册,含线上地址/curl/真实样例);README API 概览表已加两行 - -### 部署(2026-08-03 完成) - -- rsync 增量同步(**未用文档中的 --delete**,见下)→ 重启 uWSGI → 冒烟通过(列表/详情/聚合/400/404) -- 线上:`https://api.doorcome.cn/api/news/reports/`、`/api/news/events/`,Swagger `/api/docs/`「日报」tag - -### 已知事项 - -1. **服务器顶层历史平铺文件未清理**:`/home/simon/myquant/djapi/` 顶层有 views.py/urls.py/smoothBrush.py/getStockDiv2.py/env_loader.py/akshare_data.py(历史 rsync 陷阱产物),`--delete` 会删除它们,但 `divSearch.py`(离线脚本)仍绝对导入顶层 getStockDiv2/smoothBrush → 本次增量同步保留;清理前需先修 divSearch.py 的导入 -2. **既有失败测试**:`api.tests.DateFormatCorrectionTest.test_empty_string`(date_format_correction('') 期望 None 实得 ''),与本次无关 -3. 冒烟曾发现 fetch_reports 列表 SQL 缺 `r.` 别名前缀(1052 ambiguous),已修复 -4. 本地验证需绕过 macOS TCC:`HOME=/tmp/djtest_home PYTHONPATH=/tmp/djtest_pkgs`(tushare 写 ~/tk.csv 被拦 + quant 环境缺 mysql-connector-python) - -## 部署 - -```bash -# 全量同步 -rsync -avz --delete \ - --exclude='.env' --exclude='db.sqlite3' \ - --exclude='*.log' --exclude='uwsgi.pid' \ - --exclude='__pycache__/' --exclude='*.pyc' \ - --exclude='xwlb_video/' --exclude='audio_processing/' \ - /Users/summer/Downloads/cc-cursor/djapi/ \ - simon@doorcome.cn:/home/simon/myquant/djapi/ - -# 单文件同步必须写完整路径 -# 正确:rsync api/views.py simon@...:/.../djapi/api/views.py - -# 重启 -ssh simon@doorcome.cn "kill \$(lsof -ti:5004); sleep 2; /opt/miniconda/envs/django/bin/uwsgi --ini /home/simon/myquant/djapi/uwsgi.ini" -``` - -## 关键设计决策 - -- video 模块保护、向后兼容优先、不使用 python-dotenv -- rsync 陷阱:多文件源会展平路径 -- `.env` 双加载:Django 端 `djapi/env_loader.py` + video 端 `api/video/env.py` -- 股息率 TTM 用 `rolling('360D').sum()` 向量化,不手动循环 -- DeepSeek json_object 模式返回 dict,newsProcess 自适应提取 list diff --git a/djapi/docs/report_db_design.md b/djapi/docs/report_db_design.md deleted file mode 100644 index 3626b34..0000000 --- a/djapi/docs/report_db_design.md +++ /dev/null @@ -1,330 +0,0 @@ -# Milestone 10 后端实现逻辑:日报结构化入库 - -> 版本:v0.1(设计稿) | 2026-07 -> 对应 project_plan.md「十八、Milestone 10」 -> **范围**:本项目侧"后端"= 数据生产层(日报内容生成 + 结构化写入 MySQL)。 -> 不包含 API 服务与前端页面(由用户另行实现),但表结构与数据契约以本文档为准,供 API/前端对接。 - ---- - -## 1. 定位 - -现有链路:`reporter.py` 收集数据 → `_render_html()` 渲染 HTML → scp 上传 doorcome。 -改造后:`reporter.py` 收集数据 → 组装结构化 `ReportData` → 写入 MySQL(`news_report` / `news_event`),不再产出 HTML。 - -另需:把 doorcome 上 178 份历史日报 HTML(`finance_news_daily_*` ×50、`intl_news_daily_*` ×128)解析成同一 `ReportData` 结构入库。 - ---- - -## 2. 数据流总览 - -``` -[历史 HTML ×178] [每日 pipeline] - doorcome:/var/www/html/echart/research/ crawler→extractor→dedup→llm→embed→qdrant - (一次性 scp 到 data/reports_history/) │ - │ ▼ - ▼ reporter.generate_report() - report_import/parser.py │ - │ (BeautifulSoup 解析) ▼ - ▼ 组装 ReportData 组装 ReportData - report_import/importer.py │ - │ (幂等 upsert) ▼ - ▼ │ - ┌────────────────────── MySQL (myquant 库) ──────────────────────┐ - │ news_report(主表) news_event(事件明细) │ - └────────────────────────────────────────────────────────────────┘ - ▲ - API / 前端(用户另行实现,只读) -``` - ---- - -## 3. 数据模型(Pydantic,`report_db/models.py`) - -```python -class EventRow(BaseModel): - """一条事件记录,对应 news_event 一行。""" - section: str # xwlb | news | cninfo | intl - rank: int # 板块内序号(从 1 开始) - importance: int | None = None - event_type: str | None = None - title: str - summary: str | None = None - sentiment: str | None = None # positive | negative | neutral | '' - source: str | None = None # 来源(如 cls / ForexLive) - url: str | None = None - -class ReportData(BaseModel): - """一份完整日报,对应 news_report 一行 + news_event 多行。""" - report_date: date # 日报日期(YYYY-MM-DD) - report_type: str # finance | intl - file_name: str # 源文件名(新生成时可为 "") - generated_at: datetime # 生成时间 - ai_summary: str | None = None - stats: dict[str, Any] = Field(default_factory=dict) # 数据总览统计快照 → JSON 列 - events: list[EventRow] = Field(default_factory=list) -``` - ---- - -## 4. 字段映射(核心契约) - -### 4.1 事件 JSON(data/events/)→ news_event - -现有事件文件结构与 news_event 字段对应关系(reporter 收集时直接转换): - -| news_event 字段 | 事件 JSON 来源 | -| --- | --- | -| section | 来源判定:`source_id=="cninfo"` → `cninfo`;`source_id=="xwlb"` → `xwlb`;否则 `news`;intl 解析固定 `intl` | -| importance | `event.importance` | -| event_type | `event.event_type` | -| title | `title` | -| summary | `event.summary` | -| sentiment | `event.sentiment` | -| source | `source_id` | -| url | `url`(xwlb 为空) | - -### 4.2 历史 HTML → ReportData - -解析策略:**表头驱动列映射**。不同日报表格列集合不同: - -| 板块 | 表格列() | section | -| --- | --- | --- | -| 新闻联播(finance) | `# / (空) / 标题 / 重要度 / 事件类型` | xwlb | -| 重要事件:新闻(finance) | `# / (空) / 标题 / 源 / 重要度 / 事件类型 / 摘要` | news | -| 重要事件:公告调研(finance) | 同上 | cninfo | -| 重要事件(intl) | `# / (空) / 标题 / 重要度 / 事件类型 / 摘要` | intl | - -要点: - -- 以表头文本定位列索引("标题""重要度""事件类型""摘要""源"),空 `` 为情绪图标列(⚪/🔴/🟢 → neutral/negative/positive),**不要依赖列位置**。 -- 情绪图标仅存在于有图标列的表;intl 表情绪列存在,finance 表情绪列存在(空 th 首列后)。 -- intl 无"源"列时,尝试从标题尾部 `[来源]` 或摘要尾部提取,提取不到则 `source=None`。 -- 标题中的股票代码标注 `(600519, ...)` 与 ⭐(自选股标记)需剥除,只保留纯标题。 -- AI 摘要:取 `h2`("一、AI 摘要")之后紧随的 `.ai-summary` 区块纯文本(保留换行)。 -- 数据总览 → `stats` JSON:按 `h3` 标题映射 key(见 4.3),解析该 h3 后的首个 ``,缺失的板块跳过、不报错。 -- 容错:任一板块解析失败 → 记 WARNING 日志,该板块置空,不影响整份入库;整份文件解析失败 → 抛 `ReportParseError`(由 importer 捕获计数)。 - -### 4.3 数据总览 → stats JSON - -| h3 标题(含板块名) | stats key | -| --- | --- | -| M1→M6 管道 / 管道 | `pipeline`(保留原始行) | -| 各源数据 | `sources` | -| 情绪分布 | `sentiment` | -| 重要度分布 | `importance` | -| 事件类型(TOP 10 / 分布) | `event_types` | -| 文章来源分布 | `source_dist` | - -`stats` 存 MySQL `JSON` 列,前端自行解析展示。历史文件与未来新日报的 stats 结构可能不同(finance 与 intl 板块不同),一律按快照存储,不做跨版本规范化。 - ---- - -## 5. 模块设计 - -### 5.1 新包 `report_db/`(DB 层) - -``` -report_db/ -├── __init__.py # 导出 connect / init_schema / save_report -├── models.py # EventRow / ReportData(Pydantic) -├── schema.py # DDL 常量(news_report / news_event,见 project_plan.md 十八) -└── db.py # 连接、事务、写入 -``` - -`db.py` 关键函数: - -```python -def load_db_config() -> DbConfig: - """从环境变量读取 NEWS_DB_HOST/PORT/USER/PASSWORD/NAME。 - 缺失 PASSWORD 时记 ERROR 并 raise,禁止默认密码。""" - -def connect(cfg: DbConfig) -> Connection: - """pymysql.connect(autocommit=False, charset="utf8mb4", cursorclass=DictCursor)。 - 失败时 logger.exception + raise。""" - -def init_schema(conn: Connection) -> None: - """执行 schema.py 中的 CREATE TABLE IF NOT EXISTS ×2。""" - -def save_report(conn: Connection, report: ReportData) -> int: - """事务内: - 1. INSERT INTO news_report (...) VALUES (...) 或按 (report_date, report_type, file_name) - 唯一键命中时 UPDATE(新生成日报重复执行 = 覆盖同 file_name/同日期,幂等); - 2. 取 report_id,DELETE 旧事件后批量 INSERT news_event(保证整份覆盖一致)。 - 返回 report_id。""" - -def transaction(conn: Connection) -> contextmanager: - """提交/回滚上下文管理器。""" - -def fetch_report(conn: Connection, report_id: int) -> dict | None: - """读侧辅助(联调/测试用),API 侧由用户自行实现。""" -``` - -要点: - -- 所有 SQL 为 MySQL/MariaDB 方言(`JSON` 列、`ENGINE=InnoDB`、`COMMENT`),**不依赖 ORM**。 -- 连接生命周期:每次 `save_report` 短连接(report 一天跑几次,量小,无需连接池;如未来加大再换)。 -- 字符集 utf8mb4,`SET NAMES utf8mb4` 由 pymysql charset 参数处理。 - -### 5.2 新包 `report_import/`(历史解析) - -``` -report_import/ -├── __init__.py -├── parser.py # parse_finance_report / parse_intl_report(BeautifulSoup) -└── importer.py # import_history(dir, date=None, type=None) -> ImportStats -``` - -`parser.py`: - -```python -class ReportParseError(Exception): ... - -def parse_finance_report(html: str, file_name: str) -> ReportData: ... -def parse_intl_report(html: str, file_name: str) -> ReportData: ... -def parse_report(html: str, file_name: str) -> ReportData: - """按文件名前缀分流:finance_news_daily_* / intl_news_daily_*。""" -``` - -- 依赖复用现有 `beautifulsoup4`(已在 pyproject 依赖),**不新增解析库**。 -- `report_date` 从文件名解析(`*_daily_{YYYYMMDD}_*.html`),不信任目录名。 -- `generated_at` 从文件名时间(`{HHMMSS}`)或 `
` 中"生成于"文本解析,解析不到用文件 mtime。 - -`importer.py`: - -```python -@dataclass -class ImportStats: - scanned: int = 0 # 扫描到的日报文件数 - imported: int = 0 # 新入库 - skipped: int = 0 # 已存在(幂等跳过) - failed: int = 0 # 解析失败 - errors: list[str] = field(default_factory=list) - -def import_history(report_dir: Path, date: str | None = None, - report_type: str | None = None) -> ImportStats: - """遍历 {report_dir}/{YYYYMMDD}/*_news_daily_*.html, - 过滤 date / type,逐个 parse → save_report。""" -``` - -### 5.3 `scheduler/reporter.py` 改造(完全切换) - -- 新增 `_build_report_data(news, cninfo, pipeline, ai_summary, day_str, xwlb) -> ReportData`: - - 事件转换:`news["high"]` → `EventRow(section="news", ...)`;`cninfo["high"]` → `section="cninfo"`;`xwlb["items"]` → `section="xwlb"`; - - `stats` 组装:`{"pipeline": pipeline, "sources": {...}, "sentiment": news["sentiments"], "importance": news["importances"], "event_types": news["event_types"], "cninfo": {...}}`; - - 事件 `rank` 按板块内顺序编号。 -- `generate_report(day_str, *, upload=True)` 改为:收集(逻辑不变)→ `_build_report_data` → `connect()` + `save_report()`;删除 `_render_html`/`_upload` 调用。 -- `_render_*` 函数**保留但标记 deprecated**(注释说明"完全切换后不再调用"),不删除,保证最小改动、可回退。 -- 返回值由 `Path | None` 改为 `report_id: int | None`;`scheduler/pipeline.py` 中 report 步骤仅判断非 None(实施时核实该处调用,保持兼容)。 -- `stock_reporter.py` **不改动**(个股日报不在本期范围)。 - -### 5.4 `a_share_cli/main.py` 新增子命令 - -``` -uv run a-share report-import [--dir data/reports_history] [--date YYYYMMDD] [--type finance|intl] -``` - -- 默认全量扫描 `REPORT_HISTORY_DIR`(.env 可配,默认 `data/reports_history/`)。 -- 输出 ImportStats 汇总(扫描/导入/跳过/失败)。 - ---- - -## 6. 关键流程 - -### 6.1 历史导入(一次执行,可重复) - -``` -1. scp -r doorcome:/var/www/html/echart/research/2026* → data/reports_history/ - (一次手工操作,不进代码) -2. uv run a-share report-import - for each {date}/{file}: - report_type = 文件名前缀(finance|intl) - ReportData = parse_report(html, file_name) - try: save_report(conn, ReportData) → imported += 1 - except DuplicateKey: skipped += 1 # 已导入过 - except ReportParseError as e: failed += 1; errors.append(str(e)) -3. 校验: SELECT report_type, COUNT(*) FROM news_report GROUP BY report_type - 期望 50 / 128 -``` - -### 6.2 每日日报生成(pipeline 07:00 步骤) - -``` -generate_report(day_str): - news = _collect_news_events(day_str) # 不变 - cninfo = _collect_cninfo_events(day_str) # 不变 - xwlb = _collect_xwlb(day_str) # 不变 - pipeline = _collect_pipeline_stats(day_str) # 不变 - ai_summary = _generate_ai_summary(...) # 不变 - report = _build_report_data(...) # 新增 - save_report(connect(), report) # 新增(替代渲染+上传) -``` - -### 6.3 幂等策略 - -- 唯一键 `(report_date, report_type, file_name)`: - - 历史导入:命中 → 跳过(或 `--force` 覆盖); - - 新日报:`file_name=""` 时唯一键退化为 `(report_date, report_type, "")`,同一天重复跑 → UPDATE 覆盖,事件表 DELETE+INSERT 全量替换,**不产生历史残留**。 - ---- - -## 7. 配置项(.env / .env.example) - -```env -# ---- 日报结构化入库 (M10) ---- -NEWS_DB_HOST=127.0.0.1 # 开发走 ssh 隧道: ssh -L 13306:127.0.0.1:13306 pi -NEWS_DB_PORT=13306 -NEWS_DB_USER=myquant -NEWS_DB_PASSWORD= # 填真实值,禁止写入源码/文档 -NEWS_DB_NAME=myquant -REPORT_HISTORY_DIR=data/reports_history -``` - ---- - -## 8. 错误处理 - -| 场景 | 行为 | -| --- | --- | -| DB 不可达/凭据错误 | `connect()` 抛异常 → `generate_report` 记 ERROR 并返回 None(pipeline 该步骤失败,其余步骤不受影响) | -| 单份历史文件解析失败 | 记 WARNING,`failed += 1`,继续下一份;结束输出失败清单 | -| 事件字段缺失(如无摘要列) | 对应字段留 None,不抛错 | -| 全部失败 | `report-import` 返回非 0 退出码,便于排查 | - ---- - -## 9. 测试策略(tests/) - -| 文件 | 内容 | -| --- | --- | -| `tests/test_report_parser.py` | 用 fixtures(从 178 份中拷贝 finance/intl 各 1 份真实样例到 `tests/fixtures/`)断言:板块数、事件行数、字段映射、标题净化、幂等文件日期解析 | -| `tests/test_report_db.py` | 纯逻辑:`_build_report_data` 组装正确;SQL 层用 sqlite3 内存库建同构(简化 DDL)验证 upsert/覆盖语义 | -| `tests/test_report_import.py` | 临时目录构造 2-3 份假 HTML → 全流程导入 → 断言 ImportStats 计数与幂等 | -| 集成(`@pytest.mark.integration`,默认跳过) | 连真实 MySQL:init_schema + save_report + 查询回读 | - -新增 pytest marker 说明:真实 DB 连接一律走 integration,**单元测试不得依赖生产库**。 - ---- - -## 10. 依赖变更 - -- `uv add pymysql`(纯 Python 驱动,唯一新增依赖) -- 解析复用现有 `beautifulsoup4`,不新增 - ---- - -## 11. 开放问题(沿自 project_plan.md 十八,不阻塞开发) - -1. 生产连接:pi5 无法直连 `192.168.1.10:13306`(隧道仅绑 loopback)——需决定改 pi 的 autossh 绑定 / pi5 自建隧道。 -2. intl 日报生成方不在本项目,未来 intl 新日报需按同一表结构写入(本项目仅负责解析历史 + finance 新日报)。 -3. 个股日报(research 根目录文件)本期不处理。 - ---- - -## 12. 实施顺序(供开发排期) - -1. `report_db/`(models/schema/db)+ `.env` 配置 + 建表验证 -2. `report_import/parser.py` + fixtures + 单测 -3. `report_import/importer.py` + CLI `report-import` + 178 份全量导入验收 -4. `reporter.py` 改造(_build_report_data + save_report)+ pipeline 兼容性验证 -5. docs/db_schema.md 定稿(给 API/前端)、README / continuation.md 更新 diff --git a/CLAUDE.md b/docs/agent-guide.md similarity index 78% rename from CLAUDE.md rename to docs/agent-guide.md index 0cc763f..7cda9ab 100644 --- a/CLAUDE.md +++ b/docs/agent-guide.md @@ -1,4 +1,4 @@ -# CLAUDE.md +# Agent 工作指南 cc-cursor — Mac Mini 单机量化研究平台。全链路:Data → Factor → Backtest → Optimize → ML → Sentiment → Agent。 @@ -6,11 +6,11 @@ cc-cursor — Mac Mini 单机量化研究平台。全链路:Data → Factor | 模块 | 详情文件 | 核心入口 | |------|---------|---------| -| 数据层 + 数据库 | `CLAUDE-data.md` | `from data.data_manager import DataManager` | -| 因子引擎 + 情绪 | `CLAUDE-factors.md` | `from factors.registry import get_factor` | -| 回测 + 优化 | `CLAUDE-backtest.md` | `from backtest.vectorbt.engine import VectorBTEngine` | -| ML 模型 | `CLAUDE-ml.md` | `from models.lightgbm.model import LightGBMModel` | -| Agent 系统 + CLI | `CLAUDE-agents.md` | `python finance/cli/agent_cli.py daily` | +| 数据层 + 数据库 | [data-layer.md](data-layer.md) | `from data.data_manager import DataManager` | +| 因子引擎 + 情绪 | [factors.md](factors.md) | `from factors.registry import get_factor` | +| 回测 + 优化 | [backtest.md](backtest.md) | `from backtest.vectorbt.engine import VectorBTEngine` | +| ML 模型 | [ml-models.md](ml-models.md) | `from models.lightgbm.model import LightGBMModel` | +| Agent 系统 + CLI | [agents.md](agents.md) | `python finance/cli/agent_cli.py daily` | | ## 工作区布局 @@ -19,7 +19,7 @@ cc-cursor — Mac Mini 单机量化研究平台。全链路:Data → Factor | `finance/` | 核心量化引擎(**代码实际位置**)。代码内 import 用顶层名 `data.*`/`factors.*` 等 — 由 CLI 把 `finance/` 加入 sys.path;文件路径为 `finance/data/xxx.py` 等 | | `djapi/` | Django API 子项目,有独立 `djapi/CLAUDE.md` | | `shared/script/` | `autossh.sh` — MariaDB SSH 隧道 | -| `docs/` | `usage.md` / `usage.html` 使用指南、`news_report_api.md` 新闻接口文档;**新建 md 一律放这里** | +| `docs/` | 项目文档目录,包含使用指南、架构说明、API 参考等;**新建 md 一律放这里** | | `finance/strategy` `portfolio` `execution` `scheduler/` | 空壳占位(仅 `__init__.py`),逻辑未落地,别误以为有实现 | | `finance/reports/` | 日报输出 `daily_YYYYMMDD.{md,html}` | @@ -36,13 +36,13 @@ bash shared/script/autossh.sh # DB SSH 隧道 (本地 13306 → 远程 3306) | 任务类型 | 先读取 | |---------|--------| -| 数据源/数据库/cache 相关 | `CLAUDE-data.md` | -| 因子/情绪/新闻相关 | `CLAUDE-factors.md` + `CLAUDE-reference.md` | -| 回测/优化/策略相关 | `CLAUDE-backtest.md` | -| ML 模型/特征工程相关 | `CLAUDE-ml.md` | -| Agent/CLI/报告相关 | `CLAUDE-agents.md` | -| 第三方库 API/参数 | `web_fetch` / `research` 查官方文档 | -| 因子名/类名/表结构速查 | `CLAUDE-reference.md` | +| 数据源/数据库/cache 相关 | [data-layer.md](data-layer.md) | +| 因子/情绪/新闻相关 | [factors.md](factors.md) + [reference.md](reference.md) | +| 回测/优化/策略相关 | [backtest.md](backtest.md) | +| ML 模型/特征工程相关 | [ml-models.md](ml-models.md) | +| Agent/CLI/报告相关 | [agents.md](agents.md) | +| 第三方库 API/参数 | 查官方文档 | +| 因子名/类名/表结构速查 | [reference.md](reference.md) | ## 多步任务规则 diff --git a/CLAUDE-agents.md b/docs/agents.md similarity index 98% rename from CLAUDE-agents.md rename to docs/agents.md index df538db..336d803 100644 --- a/CLAUDE-agents.md +++ b/docs/agents.md @@ -1,4 +1,4 @@ -# CLAUDE-agents.md — Agent 系统 + CLI +# Agent 系统 + CLI ## Agent 架构 diff --git a/docs/api.md b/docs/api.md new file mode 100644 index 0000000..c02d323 --- /dev/null +++ b/docs/api.md @@ -0,0 +1,148 @@ +# DJAPI — Django API 参考 + +Django 5.2 项目,提供 A 股金融数据 API 和新闻联播视频处理能力。部署在 Linux 服务器上,通过 uWSGI + nginx 对外服务。 + +--- + +## 快速开始 + +```bash +# 开发服务器 +python manage.py runserver 0.0.0.0:8000 + +# uWSGI 管理 +uwsgi --ini uwsgi.ini +uwsgi --reload uwsgi.pid +uwsgi --stop uwsgi.pid + +# API 文档 +# /api/docs/ - Swagger UI +# /api/redoc/ - ReDoc +# /api/schema/ - OpenAPI Schema +``` + +## 部署 + +- 服务器:`simon@doorcome.cn`,路径 `/home/simon/myquant/djapi/` +- uWSGI 监听 `127.0.0.1:5004`,nginx 反向代理 +- 虚拟环境:`/opt/miniconda/envs/django`(Python 3.10) +- 域名:`api.doorcome.cn`、`echart.doorcome.cn` + +## 架构 + +``` +djapi/ +├── djapi/ # 项目配置 +│ ├── settings.py # Django 设置、CORS、DRF +│ ├── urls.py # 根路由 + OpenAPI schema +│ └── env_loader.py # .env 加载器 +├── api/ # 唯一 app +│ ├── views.py # 视图层(薄转发) +│ ├── urls.py # /api/* 路由 +│ ├── serializers.py # DRF Serializer(13 个) +│ ├── stock/ # 股票数据模块 +│ │ ├── data_source.py # 统一数据入口(全局单例) +│ │ ├── stock_utils.py # 通用工具 +│ │ ├── stock_basic.py # 日线行情、基本信息 +│ │ ├── getStockParam.py # 个股参数 +│ │ ├── getStockEp.py # TTM / 季度 EPS +│ │ ├── getIndexs.py # 指数行情 +│ │ ├── stockMargin.py # 融资融券 +│ │ ├── getStockFina.py # 财务报表分析 +│ │ ├── getStockDiv2.py # 股息率计算 +│ │ └── xwlbDaily.py # 新闻联播数据 +│ ├── video/ # 新闻联播视频处理(独立模块) +│ └── report/ # 日报查询 API +│ ├── query.py # 数据库查询 +│ ├── views.py # 2 个视图 +│ └── serializers.py # OpenAPI 文档 +└── uwsgi.ini # uWSGI 配置 +``` + +## 所有 API 端点(16 个) + +基础 URL:`/api/` + +| 端点 | 参数 | 说明 | +|------|------|------| +| `stockbasic/` | tscode, start_date, end_date | 日线行情 | +| `stockinfo/` | tscode | 个股基本信息 | +| `stockparam/` | tscode, start_date, end_date | 个股参数(市值等) | +| `industrys/` | industry | 按行业查股票列表 | +| `indexByName/` | index_name | 按名称查指数 | +| `indexDatas/` | tscode, start_date, end_date | 指数日行情 | +| `stockep/` | tscode, start_date, end_date | TTM EPS | +| `quarterlyEps/` | tscode, start_date, end_date | 季度 EPS | +| `finance/` | tscode, start_date, end_date | 财务报表分析 | +| `getdiv/` | tscode, start_date, end_date | 股息率(含 TTM) | +| `dailymargin/` | trade_date, exchange_id | 每日融资融券汇总 | +| `stockmargin/` | tscode, start_date, end_date | 个股融资融券 | +| `xwlbNews/` | start_date, end_date | 新闻联播(原始文本) | +| `xwlbFine/` | start_date, end_date | 新闻联播(AI 分割后) | +| `news/reports/` | report_type, start_date, end_date, id | 日报查询 | +| `news/events/` | days, importance, report_type, section, limit | 重要事件聚合 | + +### 数据源 + +所有股票数据端点通过 `api/stock/data_source.py` 统一入口: +- `get_tushare_pro()` — 全局单例(线程安全) +- `get_daily()` — 双源 fallback (Tushare → AkShare) +- `get_mysql_db()` — MySQL 全局单例 + +## 日报查询 API + +### `GET /api/news/reports/` + +| 参数 | 类型 | 默认 | 说明 | +| --- | --- | --- | --- | +| `report_type` | string | 两者 | `finance`(A 股)/ `intl`(国际) | +| `start_date` | string | 24h 前 | `YYYY-MM-DD` | +| `end_date` | string | 今天 | `YYYY-MM-DD` | +| `id` | int | 无 | 指定 id 返回单份详情(含事件) | + +### `GET /api/news/events/` + +| 参数 | 类型 | 默认 | 范围 | 说明 | +| --- | --- | --- | --- | --- | +| `days` | int | 7 | 1~365 | 最近 N 天 | +| `importance` | int | 4 | 1~5 | 最低重要度 | +| `report_type` | string | 两者 | `finance`/`intl` | 日报类型过滤 | +| `section` | string | 全部 | `xwlb`/`news`/`cninfo`/`intl` | 板块过滤 | +| `limit` | int | 100 | 1~500 | 返回条数上限 | + +详细说明见 [news_report_api.md](news_report_api.md)。 + +## 新闻联播视频处理 + +离线批处理流水线:抓取视频 → 下载 → 提取音频 → ASR 转文字 → AI 分割+取标题 → 入库。 + +```bash +python api/video/main.py # 定时任务入口 +python api/video/main_videos.py # 批量补缺 +``` + +### 模块文件 + +| 文件 | 职责 | +|------|------| +| `getVideo5.py` | 主流程:抓取→下载→ASR→入库 | +| `audioRead.py` | 音频转换、分割、ASR 识别、文本纠错 | +| `deepseek.py` | DeepSeek API 封装 | +| `newsProcess.py` | AI 新闻分割+标题提取 | +| `newsRedo.py` | 手动重处理 | +| `main.py` | 定时任务入口 | + +## 部署 + +详见 [部署说明](deployment.md)。 + +## 文档索引 + +| 文档 | 内容 | +|------|------| +| [使用指南](usage.md) | 各模块使用方法和代码示例 | +| [架构说明](architecture.md) | 项目架构、数据流、设计原则 | +| [开发指南](development.md) | 环境搭建、开发约定 | +| [部署说明](deployment.md) | 本地环境、服务器、uWSGI、rsync 部署 | +| [日报查询 API](news_report_api.md) | news/reports + news/events 详细说明 | +| [日报数据库](db_schema_v1.1.md) | news_report / news_event 表结构 | \ No newline at end of file diff --git a/docs/architecture.md b/docs/architecture.md new file mode 100644 index 0000000..257ee6d --- /dev/null +++ b/docs/architecture.md @@ -0,0 +1,196 @@ +# cc-cursor 项目架构 + +## 概述 + +cc-cursor 是一个 Mac Mini 单机量化研究平台,覆盖从数据获取到策略报告的全链路量化研究流程。 + +## 项目布局 + +``` +cc-cursor/ # 项目根目录 +│ +├── finance/ # 🔥 核心量化引擎(代码实际位置) +│ ├── config/ # 全局配置:.env 加载、数据库/API 路径 +│ ├── database/ # 数据库层:ORM 模型、SQLAlchemy 连接、DAO +│ ├── data/ # 数据层:DataManager 统一入口 +│ │ └── sources/ # Tushare + AkShare 双数据源实现 +│ ├── factors/ # 因子引擎:34 个注册因子 +│ │ ├── technical/ # 技术因子(动量、RSI、MACD、布林等 10 类) +│ │ ├── fundamental/ # 基本面因子(ROE、PE、PB、EP) +│ │ └── sentiment/ # 情绪因子(Qwen NLP + 三源新闻聚合) +│ ├── backtest/ # 回测引擎:VectorBT 封装 +│ │ ├── vectorbt/ # VectorBT 引擎适配 +│ │ └── strategies/ # 5 个内置策略(均线、RSI、动量等) +│ ├── optimizer/ # 参数优化:Optuna 引擎 + Walk-Forward +│ ├── models/ # ML 模型:LightGBM + CatBoost + 特征工程 +│ │ ├── lightgbm/ # LightGBM 模型封装 +│ │ └── catboost/ # CatBoost 模型封装 +│ ├── agents/ # Agent 系统:4 个 Agent + 编排器 +│ ├── cli/ # 命令行入口:agent_cli + 7 个 demo 验证脚本 +│ ├── reports/ # 日报输出:daily_YYYYMMDD.md + 存储层 +│ ├── strategy/ # 策略层(空壳占位,仅 __init__.py) +│ ├── portfolio/ # 组合管理(空壳占位,仅 __init__.py) +│ ├── execution/ # 执行层(空壳占位,仅 __init__.py) +│ └── scheduler/ # 调度层(空壳占位,仅 __init__.py) +│ +├── djapi/ # 🌐 Django API 后端 +│ ├── djapi/ # 项目配置:settings、urls、env_loader +│ └── api/ # 唯一 Django app +│ ├── stock/ # A 股数据 API(16 端点) +│ ├── video/ # 新闻联播视频处理(独立模块) +│ └── report/ # 日报查询 API(news/reports + news/events) +│ +├── shared/ # 🔧 共享工具 +│ └── script/ # autossh.sh(MariaDB SSH 隧道) +│ +├── docs/ # 📚 项目文档(13 个 md 文件) +│ +├── .claude/ # Claude 配置(空目录,仅保留框架) +│ +├── .git/ # Git 仓库 +│ +├── README.md # 项目入口文档 +└── .gitignore # Git 忽略规则 +``` + +### 各目录职责 + +| 目录 | 职责 | 状态 | +|------|------|------| +| `finance/` | **核心量化引擎**,全链路代码所在地 | ✅ 已实现 | +| `finance/config/` | 全局配置,`.env` 加载和路径管理 | ✅ | +| `finance/database/` | MariaDB 连接、ORM 模型、DAO 数据访问 | ✅ | +| `finance/data/` | 统一数据层,双源(Tushare→AkShare)fallback | ✅ | +| `finance/factors/` | 因子引擎,34 因子/12 分类(技术+基本面+情绪) | ✅ | +| `finance/backtest/` | VectorBT 回测,5 策略 + 截面回测 | ✅ | +| `finance/optimizer/` | Optuna 参数寻优 + Walk-Forward 验证 | ✅ | +| `finance/models/` | LightGBM/CatBoost ML 模型 + 特征工程 | ✅ | +| `finance/agents/` | 4 Agent(Research/Selection/Risk/Report)+ 编排器 | ✅ | +| `finance/cli/` | 命令行入口 + 7 个 demo 验证脚本 | ✅ | +| `finance/reports/` | 日报输出(daily_YYYYMMDD.md)+ 持久化存储 | ✅ | +| `finance/strategy/` | 策略层,预留扩展 | 🚧 空壳 | +| `finance/portfolio/` | 组合管理,预留扩展 | 🚧 空壳 | +| `finance/execution/` | 执行层,预留扩展 | 🚧 空壳 | +| `finance/scheduler/` | 调度层,预留扩展 | 🚧 空壳 | +| `djapi/` | Django API 后端,A 股数据 + 新闻联播 + 日报查询 | ✅ 已部署 | +| `shared/` | 跨项目共享工具(SSH 隧道脚本) | ✅ | +| `docs/` | 项目文档,13 个 md 文件 | ✅ | + +## 数据流 + +``` +Agent 编排层 + ├── ResearchAgent ── 因子发现(IC/IC_IR 评估) + ├── SelectionAgent ─ 多因子打分 + ML 预测 + ├── RiskAgent ────── 仓位控制 + 风险预警 + └── ReportAgent ──── 自动日报生成 + +基础引擎层 + DataManager ──→ FactorEngine ──→ BaseStrategy ──→ VectorBTEngine ──→ BacktestReport + │ │ │ + │ FeatureEngine OptunaEngine + │ │ │ + └──────→ LightGBM/CatBoost ←────────┘ + +情绪增强层 + NewsSource(AkShare/DB/MCP) ──→ QwenClient ──→ SentimentFactor ──→ FactorEngine +``` + +### 核心数据流 + +``` +Data → Factor → Model → Strategy → Backtest → Report +``` + +## 设计原则 + +### 模块隔离 +各引擎通过统一接口交互,可替换实现(VectorBT → Backtrader)。 + +### 接口标准化 +- 因子:`calculate(df) → pd.Series` +- 策略:`generate_signals(df) → pd.Series` +- 模型:`fit/predict/save/load` +- 优化:`optimize() → OptimizationResult` + +### 数据层统一 +策略/模型不直连数据源,全部通过 `DataManager`。禁止: +- 策略直接访问 AkShare/Tushare +- 模型直接访问数据库 + +### Agent 不重建轮子 +Agent 通过依赖注入复用已有引擎,编排而非重建。 + +### 防前视偏差 +- 时间序列交叉验证(TimeSeriesSplit) +- expanding window 统计量 +- 特征工程 fit 在训练集,transform 在测试集 + +## 技术栈 + +| 组件 | 技术 | 说明 | +|------|------|------| +| 数据获取 | AkShare + Tushare (双源) | Tushare 优先,AkShare fallback | +| 数据库 | MariaDB (SSH 隧道) | 本地 13306 → 远程 3306 | +| 因子/特征 | pandas / numpy / sklearn | — | +| 回测引擎 | VectorBT 1.0 | 只做多,10 万/万三 | +| 参数优化 | Optuna 4.9 | Walk-Forward 验证 | +| ML 模型 | LightGBM 4.6 + CatBoost 1.2 | 统一接口 | +| NLP 情绪 | Qwen (DashScope / Ollama) | 双后端 | +| Agent 编排 | 自研编排器 | finance/agents/ | +| API 后端 | Django 5.2 + uWSGI | djapi/ | + +## 开发进度 + +| Sprint | 模块 | 状态 | +|--------|------|------| +| Sprint 0 | 基础设施(DataManager + MariaDB 3 表) | ✅ | +| Sprint 1 | 因子引擎(34 因子 / 12 分类) | ✅ | +| Sprint 2 | VectorBT 回测(5 策略 + 截面 + BacktestReport) | ✅ | +| Sprint 3 | Optuna 优化(+ Walk-Forward) | ✅ | +| Sprint 4 | ML 模型(LightGBM + CatBoost + MLStrategy) | ✅ | +| Sprint 5 | Qwen 情绪因子(三源新闻 + 日期对齐) | ✅ | +| Sprint 6 | Agent 系统(4 Agent + CLI + 日报 .md/.html) | ✅ | +| Sprint 7 | djapi API(日报查询 ×2) | ✅ | + +**全部 8 个 Sprint 已完成。** + +## 数据源架构 + +### finance/ 引擎层 +``` +DataManager +├── TushareSource(优先,需 TUSHARE_TOKEN) +├── AkShareSource(fallback,无需 token) +└── Database Cache(SQLAlchemy + MariaDB) +``` + +### djapi/ API 层 +``` +api/stock/data_source.py(统一入口) +├── get_tushare_pro() — 全局单例 +├── get_daily() — 双源 fallback +└── get_mysql_db() — MySQL 全局单例 +``` + +## 指数代码规则 + +- `.SH` 结尾且不以 `399` 开头 → 指数(如 `000001.SH`) +- `.SZ` 开头非 `399` → 个股(如 `000001.SZ`) +- `399*.SZ` → 指数(如 `399001.SZ`) + +## 文档索引 + +| 文档 | 内容 | +|------|------| +| [使用指南](usage.md) | 各模块使用方法和代码示例 | +| [开发指南](development.md) | 环境搭建、开发约定、模块说明 | +| [因子与表结构速查](reference.md) | 34 因子注册表、DB 表结构、数据源接口 | +| [数据层详解](data-layer.md) | DataManager、数据库、缓存策略、已知 Bug | +| [因子引擎详解](factors.md) | 因子计算、情绪引擎、新闻源 | +| [回测引擎详解](backtest.md) | VectorBT、策略、信号工具、Optuna | +| [ML 模型详解](ml-models.md) | 特征工程、LightGBM/CatBoost、ML 策略 | +| [Agent 系统详解](agents.md) | Agent 架构、CLI、日报 | +| [部署说明](deployment.md) | 本地环境、服务器、uWSGI、rsync 部署 | +| [DJAPI 接口](api.md) | Django API 端点参考 | +| [日报查询 API](news_report_api.md) | news/reports + news/events 接口 | \ No newline at end of file diff --git a/CLAUDE-backtest.md b/docs/backtest.md similarity index 97% rename from CLAUDE-backtest.md rename to docs/backtest.md index e62228c..c36a133 100644 --- a/CLAUDE-backtest.md +++ b/docs/backtest.md @@ -1,4 +1,4 @@ -# CLAUDE-backtest.md — 回测引擎 + 参数优化 +# 回测引擎 + 参数优化 ## VectorBTEngine (`finance/backtest/vectorbt/engine.py`) diff --git a/CLAUDE-data.md b/docs/data-layer.md similarity index 98% rename from CLAUDE-data.md rename to docs/data-layer.md index acbf014..e5f1273 100644 --- a/CLAUDE-data.md +++ b/docs/data-layer.md @@ -1,4 +1,4 @@ -# CLAUDE-data.md — 数据层 + 数据库 +# 数据层 + 数据库 ## DataManager (`finance/data/data_manager.py`) diff --git a/djapi/docs/db_schema.md b/docs/db_schema_v1.1.md similarity index 67% rename from djapi/docs/db_schema.md rename to docs/db_schema_v1.1.md index 45bc28b..61ded50 100644 --- a/djapi/docs/db_schema.md +++ b/docs/db_schema_v1.1.md @@ -1,6 +1,7 @@ # 日报结构化入库:数据库表结构与数据契约 -> 版本:v1.0 | 2026-08-03 +> 版本:v1.1 | 2026-08-06 +> 变更(v1.0 → v1.1):新增 §3.1 stats 口径说明,修正 §3 中 pipeline "M1→M6" 的过时描述(实际 key 为 raw_total/raw_by_source/proc 等)。 > 用途:供 API / 前端对接读取日报数据。表位于 MySQL `myquant` 库,表前缀 `news_`。 > 连接:`192.168.1.10:13306`(pi 上 autossh 隧道 → doorcome.cn:3306 MariaDB 10.11),用户 `myquant`(密码在服务器 `.env` 的 `NEWS_DB_PASSWORD`)。 @@ -37,6 +38,7 @@ | summary | TEXT NULL | 摘要/正文 | | sentiment | VARCHAR(8) NULL | `positive` / `negative` / `neutral` | | source | VARCHAR(64) NULL | 来源(如 `cls`、`investinglive.com`) | +| sources | TEXT NULL | **该新闻全部来源**,JSON 数组字符串(如 `["yicai","stcn"]`);无多源/历史数据可为 `null` | | url | VARCHAR(512) NULL | 原文链接(新闻联播为空) | | created_at | DATETIME | 入库时间 | @@ -59,7 +61,7 @@ | key | finance | intl | 内容 | | --- | --- | --- | --- | -| `pipeline` | ✅ | ✅ | M1→M6 管道各环节数量:`{label: 数量}` | +| `pipeline` | ✅ | ✅ | 管道各环节数量。**finance 与 intl 内部 key 集不同**:finance 为 `{raw_total, raw_by_source, raw_total_24h, raw_by_source_24h, proc, deduped, dups, emb_count, qdrant_count, cninfo_raw}`;intl 为 `{raw_total, processed, deduped, embedded, qdrant}`(口径见 §3.1) | | `sources` | ✅ | — | 各新闻源文章数:`{源名: 数量}` | | `news` | ✅ | — | 新闻统计:`{total, hi_threshold, sentiments, importances, event_types}` | | `cninfo` | ✅ | — | 公告调研统计:`{total, hi_threshold, by_day, announcement, research, irm}` | @@ -71,6 +73,22 @@ > 历史文件与新生成日报的 stats 结构存在差异(历史为 HTML 解析快照,新生成为结构化组装),前端建议按 key 防御性读取。 +### 3.1 口径说明(重要,避免误解) + +`stats` 内各数字口径不同,请勿直接互相比较: + +| 字段 | 口径 | +| --- | --- | +| `pipeline.raw_total` | **日报日期当天**抓取的文章数(`data/raw/{src}/{date}/index.jsonl` 中 `stage=article 且 success` 的条目)。`raw_by_source` 是各源明细,**其和 = raw_total**;当天未抓取/无文章的源显示 0 | +| `pipeline.raw_total_24h` | 最近 24 小时内**抓取**(按 `fetched_at`)的文章数;`raw_by_source_24h` 为各源明细,和 = raw_total_24h。当天 07:00 抓取的数据其值 ≈ raw_total(并非"24h 内发布的新闻",raw 层无发布时间的可靠字段) | +| `pipeline.proc / deduped / dups / emb_count / qdrant_count` | 抽取 / 去重后 / 重复 / 向量化 / Qdrant 总量(`qdrant_count` 为全量累计,非当天) | +| `news.total` | **过去 30 小时窗口内**经 LLM 抽取的新闻事件数。**≠ raw_total**:raw 是抓取的文章数,news 是抽取后的事件数(会有过滤/合并),两者不可互相验证 | +| `news.importances` | `{重要度等级(1-5): 事件数}`,**各等级之和 = news.total** | +| `news.sentiments` | `{情绪: 事件数}`(positive/negative/neutral),和 = news.total | +| `news.event_types` | `{事件类型: 事件数}`(TOP 10) | +| `pipeline.*(finance)` | finance 日报专用:`proc`=抽取后条数、`deduped`=去重后、`dups`=重复条数、`emb_count`=向量化条数、`qdrant_count`=Qdrant 全量累计(非当天)、`cninfo_raw`=当天 cninfo 抓取数。`raw_by_source` 各源之和 = `raw_total` | +| `pipeline.*(intl)` | intl 日报**结构不同**:`raw_total`=当天抓取文章数、`processed`=处理数(≈ raw_total)、`deduped`=去重后条数、`embedded`=向量化条数、`qdrant`=Qdrant 全量累计。intl 无 `raw_by_source` 明细与 `cninfo_raw` | + --- ## 4. 常用查询示例(API 实现参考) diff --git a/docs/deployment.md b/docs/deployment.md new file mode 100644 index 0000000..a53ea6d --- /dev/null +++ b/docs/deployment.md @@ -0,0 +1,274 @@ +# 部署说明 + +cc-cursor 包含两个运行组件:**finance 量化引擎**(Mac Mini 本地)和 **djapi API 后端**(Linux 服务器)。部署架构如下: + +``` +┌─ Mac Mini (本地) ─────────────────────────────────────────────┐ +│ │ +│ finance/ 量化引擎 │ +│ ├── 数据获取 (AkShare/Tushare) │ +│ ├── 因子计算 / 回测 / ML │ +│ └── Agent 日报生成 │ +│ │ +│ shared/script/autossh.sh │ +│ └── SSH 隧道 :13306 ──────────────────────┐ │ +│ │ │ +└─────────────────────────────────────────────┼──────────────────┘ + │ + MariaDB 10.11 │ + doorcome.cn:3306│ + │ +┌─ Linux 服务器 (doorcome.cn) ────────────────┼──────────────────┐ +│ │ │ +│ djapi/ Django API │ │ +│ ├── uWSGI :5004 │ │ +│ ├── nginx 反向代理 │ │ +│ └── 域名: api.doorcome.cn │ │ +│ │ │ +│ MariaDB myquant 库 │ │ +│ ├── mac_* 表 (量化引擎数据) │ │ +│ ├── xwlb_* 表 (新闻联播) │ │ +│ └── news_* 表 (日报) │ │ +│ │ │ +└─────────────────────────────────────────────┘ │ +``` + +--- + +## 1. 本地环境(Mac Mini) + +### 1.1 Python 环境 + +```bash +conda activate quant # Python 3.11.13 +``` + +### 1.2 环境变量 + +复制并编辑 `finance/.env`(参考 `finance/.env.example`): + +```bash +TUSHARE_TOKEN=your_token_here +QWEN_API_KEY=sk-your-key-here +MAC_DB_HOST=127.0.0.1 +MAC_DB_PORT=13306 +MAC_DB_USER=myquant +MAC_DB_PASSWORD=your_password_here +MAC_DB_NAME=myquant +``` + +### 1.3 数据库 SSH 隧道 + +```bash +bash shared/script/autossh.sh +``` + +脚本内容: + +```bash +autossh -M 0 -fN -L 13306:localhost:3306 tunnel@doorcome.cn +``` + +验证隧道: + +```bash +lsof -i :13306 | grep LISTEN +``` + +连接信息: + +``` +Host: 127.0.0.1 +Port: 13306 +User: myquant +Database: myquant +``` + +--- + +## 2. 服务器环境(doorcome.cn) + +### 2.1 基本信息 + +| 项目 | 值 | +|------|-----| +| 服务器 | `simon@doorcome.cn` | +| 项目路径 | `/home/simon/myquant/djapi/` | +| Python 环境 | `/opt/miniconda/envs/django` (Python 3.10) | +| uWSGI 端口 | `127.0.0.1:5004` | +| nginx 反向代理 | 域名 → `127.0.0.1:5004` | +| 生产域名 | `api.doorcome.cn`、`echart.doorcome.cn` | + +### 2.2 服务器环境变量 + +服务器端 `djapi/.env`(**不随代码同步**,需在服务器上手动维护): + +```bash +DJANGO_SECRET_KEY=... +TUSHARE_TS_TOKEN=... +MYSQL_HOST=localhost +MYSQL_PORT=3306 +MYSQL_USER=myquant +MYSQL_PASSWORD=... +MYSQL_DATABASE=myquant +NEWS_DB_HOST=127.0.0.1 +NEWS_DB_PORT=3306 +NEWS_DB_USER=myquant +NEWS_DB_PASSWORD=... +NEWS_DB_NAME=myquant +DEEPSEEK_API_KEY=... +DASHSCOPE_API_KEY=... +``` + +### 2.3 uWSGI 配置 + +配置文件:`djapi/uwsgi.ini` + +```ini +[uwsgi] +http = 127.0.0.1:5004 +chdir = /home/simon/myquant/djapi +module = djapi.wsgi:application +uid = simon +gid = simon +master = true +workers = 5 +pidfile = /home/simon/myquant/djapi/uwsgi.pid +vacuum = true +thunder-lock = true +enable-threads = true +harakiri = 30 +post-buffering = 4096 +daemonize = /home/simon/myquant/djapi/uwsgi.log +log-maxsize = 10240000 +py-autoreload = 1 +virtualenv = /opt/miniconda/envs/django +env = DJANGO_SETTINGS_MODULE=djapi.settings +``` + +### 2.4 uWSGI 管理命令 + +```bash +# 启动 +uwsgi --ini uwsgi.ini + +# 热重载 +uwsgi --reload uwsgi.pid + +# 停止 +uwsgi --stop uwsgi.pid + +# 强制停止 +kill $(lsof -ti:5004) +``` + +--- + +## 3. 代码部署 + +### 3.1 全量同步 + +```bash +rsync -avz --delete \ + --exclude='.env' --exclude='db.sqlite3' \ + --exclude='*.log' --exclude='uwsgi.pid' \ + --exclude='__pycache__/' --exclude='*.pyc' \ + --exclude='xwlb_video/' --exclude='audio_processing/' \ + /path/to/cc-cursor/djapi/ \ + simon@doorcome.cn:/home/simon/myquant/djapi/ +``` + +### 3.2 单文件同步 + +```bash +# 必须写完整目标路径,否则会展平到根目录 +rsync -avz api/views.py simon@doorcome.cn:/home/simon/myquant/djapi/api/views.py +``` + +### 3.3 重启服务 + +```bash +ssh simon@doorcome.cn "kill \$(lsof -ti:5004); sleep 2; /opt/miniconda/envs/django/bin/uwsgi --ini /home/simon/myquant/djapi/uwsgi.ini" +``` + +### 3.4 部署后验证 + +```bash +# Swagger 文档 +curl -s https://api.doorcome.cn/api/docs/ | head -5 + +# 日报查询 +curl -s 'https://api.doorcome.cn/api/news/reports/' | python -m json.tool | head -20 + +# 健康检查(冒烟) +curl -s -o /dev/null -w "%{http_code}" 'https://api.doorcome.cn/api/news/reports/' +# → 200 +``` + +--- + +## 4. 数据库 + +### 4.1 表前缀 + +| 前缀 | 用途 | 位置 | +|------|------|------| +| `mac_` | 量化引擎数据(股票列表、日线、财务、报告) | finance 引擎写入 | +| `xwlb_` | 新闻联播数据 | djapi video 模块写入 | +| `news_` | 日报数据 | 外部 pipeline 写入,djapi 只读 | + +### 4.2 量化引擎表 + +| 表 | 内容 | 主键 | +|----|------|------| +| `mac_stock_basic` | A 股列表 (5,524 只) | ts_code | +| `mac_stock_daily` | 日线 OHLCV | (ts_code, trade_date) | +| `mac_stock_financial` | 财务指标 | (ts_code, end_date) | +| `mac_report` | 报告持久化 | id | + +### 4.3 日报表 + +| 表 | 内容 | +|----|------| +| `news_report` | 日报主表(一行 = 一份日报) | +| `news_event` | 日报事件明细(一行 = 一条事件) | + +详见 [db_schema_v1.1.md](db_schema_v1.1.md)。 + +--- + +## 5. 开发环境 + +### 5.1 本地运行 djapi + +```bash +cd djapi +python manage.py runserver 0.0.0.0:8000 +``` + +### 5.2 API 文档 + +- Swagger UI:`/api/docs/` +- ReDoc:`/api/redoc/` +- OpenAPI Schema:`/api/schema/` + +### 5.3 数据库初始化 + +```bash +python manage.py makemigrations +python manage.py migrate +python manage.py createsuperuser +``` + +--- + +## 6. 文档索引 + +| 文档 | 内容 | +|------|------| +| [使用指南](usage.md) | 各模块使用方法和代码示例 | +| [架构说明](architecture.md) | 项目架构、数据流、设计原则 | +| [开发指南](development.md) | 环境搭建、开发约定 | +| [DJAPI 接口](api.md) | Django API 端点参考 | +| [日报查询 API](news_report_api.md) | news/reports + news/events 接口 | +| [日报数据库](db_schema_v1.1.md) | news_report / news_event 表结构 | \ No newline at end of file diff --git a/docs/development.md b/docs/development.md new file mode 100644 index 0000000..2aab14e --- /dev/null +++ b/docs/development.md @@ -0,0 +1,271 @@ +# cc-cursor 开发指南 + +## 环境搭建 + +### 1. 克隆项目 + +```bash +git clone https://github.com/Simon2046/myquant.git +cd cc-cursor +``` + +### 2. Python 环境 + +```bash +conda activate quant # Python 3.11.13 +``` + +### 3. 数据库 SSH 隧道 + +```bash +bash shared/script/autossh.sh +# host: 127.0.0.1:13306 user: myquant database: myquant +``` + +### 4. 环境变量配置 + +复制并编辑 `finance/.env`(参考 `finance/.env.example`): + +```bash +TUSHARE_TOKEN=your_token_here +QWEN_API_KEY=sk-your-key-here +MAC_DB_PASSWORD=your_password_here +``` + +### 5. 验证环境 + +```bash +python finance/cli/demo_data_manager.py +``` + +## 项目结构 + +``` +finance/ # 核心量化引擎(代码实际位置) +├── config/ # 全局配置 +│ └── settings.py # .env 加载 + 路径配置 +├── database/ # 数据库层 +│ ├── connection.py # SQLAlchemy 引擎 + SSH 自动恢复 +│ ├── models.py # ORM 模型(mac_ 前缀表) +│ └── dao.py # 数据访问对象 +├── data/ # 数据层 +│ ├── data_manager.py # 统一数据入口 +│ └── sources/ # 数据源实现 +│ ├── tushare_source.py +│ └── akshare_source.py +├── factors/ # 因子引擎 +│ ├── base.py # 因子基类 +│ ├── engine.py # 因子计算引擎 +│ ├── registry.py # 因子注册表 +│ ├── technical/ # 技术因子(10 类) +│ ├── fundamental/ # 基本面因子 +│ └── sentiment/ # 情绪因子 +│ ├── sentiment_engine.py +│ ├── sentiment_factor.py +│ ├── news_source.py # 三源新闻聚合 +│ └── qwen_client.py # Qwen API 客户端 +├── backtest/ # 回测引擎 +│ ├── base.py # 策略基类 +│ ├── report.py # 回测报告 +│ ├── signal.py # 信号工具 +│ ├── vectorbt/ # VectorBT 引擎 +│ └── strategies/ # 内置策略 +├── optimizer/ # 参数优化 +│ ├── engine.py # Optuna 引擎 +│ ├── space.py # 搜索空间 +│ ├── objectives.py # 优化目标 +│ └── result.py # 优化结果 +├── models/ # ML 模型 +│ ├── base.py # 模型基类 +│ ├── features.py # 特征工程 +│ ├── backtest_integration.py # ML 策略 +│ ├── lightgbm/ +│ └── catboost/ +├── agents/ # Agent 系统 +│ ├── base.py # Agent 基类 +│ ├── orchestrator.py # 编排器 +│ ├── research_agent.py # 因子研究 +│ ├── selection_agent.py # 股票打分 +│ ├── risk_agent.py # 风险评估 +│ └── report_agent.py # 日报生成 +├── cli/ # 命令行 +│ ├── agent_cli.py # Agent CLI 入口 +│ └── demo_*.py # 验证脚本 +└── reports/ # 日报输出 + └── storage.py # 报告持久化 +``` + +## 开发约定 + +### 代码组织 + +- 代码内 import 用顶层名 `data.*`/`factors.*` 等 — CLI 自动把 `finance/` 加入 sys.path +- 文件路径为 `finance/data/xxx.py` 等 + +### 数据流约束 + +``` +Data → Factor → Model → Strategy → Backtest → Report +``` + +**必须遵守**: +- 策略层禁止直接访问 AkShare/Tushare → 全部通过 `DataManager` +- 模型层禁止直接访问数据库 → 全部通过 `DataManager` +- 指数代码规则:`.SH`=指数, `.SZ` 开头非 399=个股 +- 数据源优先级:Tushare → AkShare (fallback) + +### 接口规范 + +#### 因子接口 +```python +class BaseFactor: + def calculate(self, df: pd.DataFrame) -> pd.Series: + """接收 OHLCV 数据,返回因子值序列""" +``` + +#### 策略接口 +```python +class BaseStrategy: + def generate_signals(self, factor_df: pd.DataFrame) -> pd.Series: + """接收因子数据,返回交易信号: 1=buy, 0=sell, -1=hold""" +``` + +#### 模型接口 +```python +class BaseModel: + def fit(self, X, y): ... + def predict(self, X) -> np.ndarray: ... + def save(self, path): ... + @classmethod + def load(cls, path): ... +``` + +### 多步任务规则 + +复杂任务(涉及 3+ 文件或 2+ 模块)执行前: +1. 输出执行计划清单(步骤 + 每步验证方法) +2. 每步完成后验证通过才继续 +3. 遇到失败先定位根因,不跳过 + +### 修改多文件前 + +先说明:文件清单、原因、影响;优先小范围修改。 + +## 模块说明 + +### 数据层 (`finance/data/`) + +双数据源架构:Tushare(优先)→ AkShare(fallback),DB 缓存优先。 + +```python +from data.data_manager import DataManager +dm = DataManager() +dm.init_db() # 首次建表(幂等) +stocks = dm.get_stock_list() # → 5,524 只 +daily = dm.get_daily("000001.SZ") # → 日线 +fina = dm.get_financial("000001.SZ") # → 财务 +n = dm.sync_daily("000001.SZ") # → 增量同步 +``` + +### 因子引擎 (`finance/factors/`) + +34 个注册因子,12 个分类:动量、RSI、MACD、量价、布林、ATR、均线、波动率、换手率、振幅、基本面、情绪。 + +```python +from factors.registry import get_factor, list_factors +from factors.engine import FactorEngine + +fe = FactorEngine(dm) +factor_df = fe.compute("000001.SZ", [get_factor("momentum_20"), get_factor("rsi_14")]) +``` + +### 回测引擎 (`finance/backtest/`) + +VectorBT 1.0,只做多,10 万/万三。5 个内置策略 + 自定义策略接口。 + +```python +from backtest.vectorbt.engine import VectorBTEngine +engine_bt = VectorBTEngine(initial_capital=100_000, commission=0.0003) +report = engine_bt.run(strategy, price_df, factor_df) +``` + +### 参数优化 (`finance/optimizer/`) + +Optuna 4.9 + Walk-Forward 滚动验证。 + +```python +from optimizer.engine import OptunaEngine +opt = OptunaEngine(engine_bt) +result = opt.optimize(StrategyClass, space, price_df, factor_df, metric="sharpe", n_trials=200) +``` + +### ML 模型 (`finance/models/`) + +LightGBM 4.6 + CatBoost 1.2,统一接口,特征工程防前视偏差。 + +```python +from models.features import FeatureEngine +from models.lightgbm.model import LightGBMModel + +fe = FeatureEngine(lookahead=5) +X, y = fe.build(factor_df, price_df, fit=True) +model = LightGBMModel(params={"n_estimators": 200}).fit(X_train, y_train) +``` + +### 情绪因子 (`finance/factors/sentiment/`) + +三数据源聚合(AkShare 个股新闻 + 新闻联播 DB + MCP trendradar-news),Qwen DashScope + Ollama 双后端。 + +```python +from factors.sentiment.sentiment_engine import SentimentEngine +sent = SentimentEngine(dm) +sent_df = sent.compute("000001.SZ", max_news=20) +``` + +### Agent 系统 (`finance/agents/`) + +4 个 Agent(Research/Selection/Risk/Report)+ 编排器 + CLI。 + +```bash +python finance/cli/agent_cli.py daily # 完整流程 +python finance/cli/agent_cli.py picks 15 # 选股 +python finance/cli/agent_cli.py risk # 风险评估 +``` + +## 安全规范 + +### 禁止提交 + +- `.env`(含真实 key) +- API Key(`sk-*`、`TUSHARE_TOKEN` 等) +- Cookie / Session / Token / 密钥 +- 个人隐私数据(手机号、身份证、密码) + +### 必须提供 + +- `.env.example` — 仅含占位符的示例配置 + +### 提交前检查 + +```bash +grep -r "sk-\|token\|password" --include="*.py" --include="*.md" --include="*.yaml" | grep -v ".example\|your_token\|your_password" +``` + +## 已知 Bug 速查 + +详见 [data-layer.md](data-layer.md) 和 [agents.md](agents.md) 中的 Bug 列表。 + +## 文档索引 + +| 文档 | 内容 | +|------|------| +| [使用指南](usage.md) | 各模块使用方法和代码示例 | +| [架构说明](architecture.md) | 项目架构、数据流、设计原则 | +| [因子与表结构速查](reference.md) | 34 因子注册表、DB 表结构、数据源接口 | +| [数据层详解](data-layer.md) | DataManager、数据库、缓存策略、已知 Bug | +| [因子引擎详解](factors.md) | 因子计算、情绪引擎、新闻源 | +| [回测引擎详解](backtest.md) | VectorBT、策略、信号工具、Optuna | +| [ML 模型详解](ml-models.md) | 特征工程、LightGBM/CatBoost、ML 策略 | +| [Agent 系统详解](agents.md) | Agent 架构、CLI、日报 | +| [部署说明](deployment.md) | 本地环境、服务器、uWSGI、rsync 部署 | +| [DJAPI 接口](api.md) | Django API 端点参考 | \ No newline at end of file diff --git a/CLAUDE-factors.md b/docs/factors.md similarity index 97% rename from CLAUDE-factors.md rename to docs/factors.md index ca568f2..d904ab9 100644 --- a/CLAUDE-factors.md +++ b/docs/factors.md @@ -1,4 +1,4 @@ -# CLAUDE-factors.md — 因子引擎 + 情绪因子 +# 因子引擎 + 情绪因子 ## FactorEngine (`finance/factors/engine.py`) diff --git a/CLAUDE-ml.md b/docs/ml-models.md similarity index 98% rename from CLAUDE-ml.md rename to docs/ml-models.md index 887b6a5..d7313e2 100644 --- a/CLAUDE-ml.md +++ b/docs/ml-models.md @@ -1,4 +1,4 @@ -# CLAUDE-ml.md — ML 模型层 +# ML 模型层 ## FeatureEngine (`finance/models/features.py`) diff --git a/docs/news_report_api.md b/docs/news_report_api.md index 5625c40..1f8de71 100644 --- a/docs/news_report_api.md +++ b/docs/news_report_api.md @@ -1,7 +1,7 @@ # 日报查询 API 使用手册(news_report / news_event) > 版本:v1.1 | 2026-08-03(已部署至生产 `api.doorcome.cn`) -> 数据表定义见 `djapi/docs/db_schema.md`,数据生产侧设计见 `djapi/docs/report_db_design.md`。 +> 数据表定义见 [db_schema_v1.1.md](db_schema_v1.1.md)。 > 本文档为前端/调用方对接手册:两个只读查询接口的线上地址、参数、curl 用法与返回结构。 --- @@ -61,11 +61,22 @@ Swagger 交互式文档(自动生成,含全部参数说明):`https://api "generated_at": "2026-08-03T07:00:00", "ai_summary": "……(AI 摘要全文,按条目分行)", "stats": { - "pipeline": { "M1": 210, "M2": 205, "M3": 198, "M4": 190, "M5": 188, "M6": 180 }, - "sources": { "cls": 95, "eastmoney": 60, "other": 25 }, - "news": { "total": 180, "hi_threshold": 18, "sentiments": {...}, "importances": [...], "event_types": [...] }, - "cninfo": { "total": 58, "hi_threshold": 4, "by_day": {...}, "announcement": 40, "research": 15, "irm": 3 }, - "xwlb": { "total": 28, "date": "2026-08-03" } + "pipeline": { + "raw_total": 116, + "raw_total_24h": 116, + "raw_by_source": { "经济观察网": 31, "新浪财经": 27, "证券时报": 28, "财联社": 15, "东方财富": 7, "中证券网": 2, "第一财经": 3, "中国证券网": 3 }, + "raw_by_source_24h": { ... }, + "proc": ..., "deduped": ..., "dups": ..., "emb_count": ..., "qdrant_count": ..., "cninfo_raw": ... + }, + "news": { + "total": 643, + "hi_threshold": 4, + "sentiments": { "neutral": 470, "positive": 123, "negative": 50 }, + "importances": { "1": 115, "2": 275, "3": 185, "4": 48, "5": 20 }, + "event_types": { "其他": 287, "国际局势": 107, "宏观政策": 62, "行业政策": 49, "财报披露": 29 } + }, + "cninfo": { "total": 48, "hi_threshold": 2, "by_day": { "2026-08-06": 8, "2026-08-01": 9 }, "announcement": 47, "research": 1, "irm": 0 }, + "xwlb": { "total": 21, "date": "08月05日" } }, "created_at": "2026-08-03T07:01:00" }, @@ -77,18 +88,26 @@ Swagger 交互式文档(自动生成,含全部参数说明):`https://api "generated_at": "2026-08-03T07:00:00", "ai_summary": "……", "stats": { - "pipeline": { "M1": 150, "M2": 145, "M3": 140, "M4": 135, "M5": 130, "M6": 125 }, - "sentiment": [...], - "importance": [{ "重要度": 5, "数量": 3 }, { "重要度": 4, "数量": 12 }], - "event_types": [{ "事件类型": "地缘政治", "数量": 8 }], - "source_dist": [{ "来源": "investinglive.com", "文章数": 45 }] + "pipeline": { "raw_total": 442, "processed": 442, "deduped": 111, "embedded": 112, "qdrant": 2952 }, + "importance": [ + { "importance": 2, "count": 2 }, { "importance": 3, "count": 31 }, + { "importance": 4, "count": 26 }, { "importance": 5, "count": 5 } + ], + "event_types": [ + { "event_type": "宏观经济", "count": 15 }, { "event_type": "地缘政治", "count": 12 }, + { "event_type": "财报披露", "count": 9 }, { "event_type": "央行决议", "count": 7 } + ], + "source_dist": [ + { "source": "InvestingLive", "count": 30 }, { "source": "MarketWatch", "count": 3 }, + { "source": "CNBC", "count": 3 }, { "source": "Yahoo Finance", "count": 1 } + ] }, "created_at": "2026-08-03T07:01:00" } ] ``` -> `stats` 为 JSON 快照:finance 含 `pipeline/sources/news/cninfo/xwlb`,intl 含 `pipeline/sentiment/importance/event_types/source_dist`。前端按 key 防御性读取(见 db_schema.md §3)。 +> `stats` 为 JSON 快照(示例为 2026-08-06 真实数据):finance 含 `pipeline/news/cninfo/xwlb`,intl 含 `pipeline/importance/event_types/source_dist`(另有 `sentiment`)。finance 与 intl 的 pipeline 内部 key 集**不同**(finance: `raw_total/raw_by_source/proc/dups/...`;intl: `raw_total/processed/deduped/embedded/qdrant`),前端按 key 防御性读取;各数字口径(`raw_total` ≠ `news.total` 等)见 `db_schema_v1.1.md` §3.1。 **传 `id`** → 单份详情(对象),增加 `events` 数组(按 `section, rank` 排序)。真实返回(id=182,58 条事件): @@ -179,12 +198,15 @@ curl 'https://api.doorcome.cn/api/news/reports/?id=182' "title": "特朗普称已取消对伊朗的袭击计划,因双方就协议框架达成一致", "summary": "……", "sentiment": "neutral", - "source": "investinglive.com", + "source": "InvestingLive", + "sources": ["InvestingLive"], "url": "https://www.investing.com/..." } ] ``` +> `sources` 为该新闻的**全部来源**(JSON 数组,字符串列表);`source` 为主来源(单值)。多源事件(如同一新闻被多家媒体转载)时 `sources` 含多个元素;历史事件可能为 `null`。 + ### curl 示例 ```bash diff --git a/CLAUDE-reference.md b/docs/reference.md similarity index 98% rename from CLAUDE-reference.md rename to docs/reference.md index adf9258..a5a4ef4 100644 --- a/CLAUDE-reference.md +++ b/docs/reference.md @@ -1,4 +1,4 @@ -# CLAUDE-reference.md — 因子 + 表结构速查 +# 因子 + 表结构速查 ## 因子速查(34 个,12 分类) diff --git a/docs/usage.html b/docs/usage.html deleted file mode 100644 index 59424cb..0000000 --- a/docs/usage.html +++ /dev/null @@ -1,1512 +0,0 @@ - - - - - -cc-cursor 使用指南 - - - -
-

cc-cursor 使用指南

-
- -
-

Mac Mini 单机量化研究平台,覆盖数据获取 → 因子计算 → 回测 → 参数优化 → ML 模型 → 情绪因子 → Agent 系统 全链路。

-
-
-

1. 环境准备

-

硬件要求

- -

Python 环境

-
# 激活 conda 环境
-conda activate quant
-
-# 确认 Python 版本
-python --version  # → 3.11.13
-
- -

核心依赖

-
- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -
包版本用途
pandas3.0数据处理
numpy2.4数值计算
akshare1.18A 股数据获取
vectorbt1.0回测引擎
optuna4.9参数优化
lightgbm4.6梯度提升模型
catboost1.2梯度提升模型
scikit-learn1.9特征工程
sqlalchemy2.0数据库 ORM
pymysql1.2MySQL 连接
-

项目路径设置

-

所有代码从 finance/ 目录运行。Python 脚本开头加入:

-
import sys
-sys.path.insert(0, "/path/to/cc-cursor/finance")
-
- -
-

2. 数据库连接

-

建立 SSH 隧道

-

系统通过 SSH 隧道连接远程 MariaDB:

-
bash shared/script/autossh.sh
-
- -

验证隧道:

-
lsof -i :13306 | grep LISTEN
-# → ssh ... localhost:13306 (LISTEN) ...
-
- -

连接信息

-
Host:     127.0.0.1
-Port:     13306
-User:     myquant
-Password: <your-db-password>
-Database: myquant
-
- -

数据库表

-

所有表使用 mac_ 前缀,与已有表隔离:

- - - - - - - - - - - - - - - - - - - - - - - - - -
表名内容说明
mac_stock_basicA 股列表5,524 只股票
mac_stock_daily日线行情按需同步
mac_stock_financial财务指标同花顺核心指标
-

测试连接

-
from database.connection import test_connection
-
-if test_connection():
-    print("数据库连接成功")
-else:
-    print("请先建立 SSH 隧道: bash shared/script/autossh.sh")
-
- -
-

3. 数据层 — DataManager

-

初始化

-
from data.data_manager import DataManager
-
-dm = DataManager()
-dm.init_db()  # 首次使用创建表(幂等操作)
-
- -

获取股票列表

-
# 从 DB 缓存读取(已缓存的 5,524 只 A 股)
-stocks = dm.get_stock_list()
-# → DataFrame: index=ts_code, columns=[name, area, industry, ...]
-
-# 强制从 AkShare 刷新
-stocks = dm.get_stock_list(force_refresh=True)
-
- -

获取日线数据

-
# 获取单只股票日线(DB 缓存优先,缺失自动补拉)
-daily = dm.get_daily("000001.SZ")
-# → DataFrame: trade_date, open, high, low, close, vol, amount, ...
-
-# 指定日期范围
-daily = dm.get_daily("000001.SZ", start="20240101", end="20241231")
-
-# 直接用索引
-price = daily.set_index("trade_date").sort_index()
-close = price["close"]
-
- -

获取财务数据

-
fina = dm.get_financial("000001.SZ")
-# → DataFrame: end_date, eps, bvps, roe, net_profit_margin, debt_to_assets, ...
-# 数据源: stock_financial_abstract_ths(同花顺)
-# 覆盖: 主板/创业板/科创板
-
- -

数据同步

-
# 增量同步:从 DB 最新日期到今天的缺失数据
-n = dm.sync_daily("000001.SZ")
-
-# 批量同步全部股票(谨慎使用,耗时长)
-total = dm.sync_all_daily()
-
- -
-

4. 因子引擎 — FactorEngine

-

因子注册表

-
from factors.registry import get_factor, list_factors, list_categories
-
-# 查看所有因子分类
-print(list_categories())
-# → ['动量', 'RSI', 'MACD', '量价', '布林', 'ATR', '均线', '波动率', '换手率', '振幅', '基本面', '情绪']
-
-# 查看某个分类下的因子
-print(list_factors("RSI"))
-# → ['rsi_7', 'rsi_14']
-
-# 查看全部因子
-all_factors = list_factors()
-print(len(all_factors))
-# → 34
-
- -

创建因子实例

-
# 按名称获取(使用默认参数)
-factor = get_factor("momentum_20")    # 20 日动量
-factor = get_factor("rsi_14")         # 14 日 RSI
-factor = get_factor("roe")            # ROE 基本面因子
-factor = get_factor("news_sent_5")    # 5 日新闻情绪因子
-
-# 自定义参数
-from factors.technical.momentum import MomentumFactor
-factor = MomentumFactor(period=60)
-
- -

计算因子

-
from factors.engine import FactorEngine
-
-engine_fe = FactorEngine(dm)
-
-# 单股票多因子
-factors = [
-    get_factor("momentum_20"),
-    get_factor("rsi_14"),
-    get_factor("volatility_20"),
-    get_factor("ma_dev_20"),
-]
-factor_df = engine_fe.compute("000001.SZ", factors)
-# → DataFrame: index=trade_date, columns=[momentum_20, rsi_14, volatility_20, ma_dev_20]
-
-# 查看因子值
-print(factor_df.tail())
-print(factor_df.describe())
-
- -

截面因子

-
# 计算多只股票在某一天的因子值
-cross = engine_fe.compute_universe(
-    factors=[get_factor("momentum_20"), get_factor("rsi_14")],
-    date="20250630",
-    ts_codes=["000001.SZ", "600519.SH", "300750.SZ"],
-)
-# → DataFrame: index=ts_code, columns=[momentum_20, rsi_14]
-
- -

因子质量检查

-
# 查看 NaN 率
-total = len(factor_df)
-for col in factor_df.columns:
-    nan_pct = factor_df[col].isna().sum() / total * 100
-    print(f"{col}: NaN {nan_pct:.1f}%")
-# 正常范围: 技术因子 0.3%-2.1%, 基本面因子 0%
-
- -
-

5. 回测引擎 — VectorBTEngine

-

参数

- -

创建引擎

-
from backtest.vectorbt.engine import VectorBTEngine
-
-engine_bt = VectorBTEngine(
-    initial_capital=100_000,
-    commission=0.0003,
-)
-
- -

使用内置策略

-
from backtest.strategies.rsi_mean_revert import RSIMeanRevertStrategy
-
-# 创建策略
-strategy = RSIMeanRevertStrategy(oversold=30, overbought=70)
-
-# 运行回测
-report = engine_bt.run(strategy, price_df, factor_df)
-
- -

读取回测报告

-
# 一行摘要
-print(report.summary())
-# → 收益=29.4% 年化=4.3% 回撤=-19.1% 夏普=0.37 胜率=77.1% 交易=70笔
-
-# 字典格式
-metrics = report.to_dict()
-# → {'total_return': 29.4, 'cagr': 4.3, 'sharpe_ratio': 0.37, ...}
-
-# 获取净值曲线
-equity = report.equity_curve    # pd.Series
-drawdown = report.drawdown_curve  # pd.Series
-
-# 逐笔交易
-trades = report.trades_df       # pd.DataFrame
-
- -

内置策略清单

- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -
策略类名适用场景
均线交叉SMACrossStrategy(fast=5, slow=20)趋势跟踪
RSI 反转RSIMeanRevertStrategy(oversold=30, overbought=70)均值回归
动量突破MomentumBreakoutStrategy(lookback=20, exit_period=10)动量策略
因子阈值FactorCrossStrategy(factor_column, buy_threshold, sell_threshold)通用因子
因子轮动FactorRotationStrategy(factor_name, top_n=5)截面选股
-

自定义策略

-
from backtest.base import BaseStrategy
-
-class MyStrategy(BaseStrategy):
-    name = "my_strategy"
-    category = "custom"
-
-    def __init__(self, param_a=10):
-        self.param_a = param_a
-
-    def generate_signals(self, factor_df):
-        # factor_df 包含因子值和 close 列
-        # 返回: 1=买入, 0=卖出, -1=持有
-        signals = pd.Series(-1, index=factor_df.index)
-        signals[factor_df["rsi_14"] < 30] = 1   # RSI 超卖买入
-        signals[factor_df["rsi_14"] > 70] = 0   # RSI 超买卖出
-        return signals
-
-report = engine_bt.run(MyStrategy(param_a=20), price_df, factor_df)
-
- -

截面回测(多股票)

-
report_xs = engine_bt.run_cross_section(
-    strategy,
-    price_universe={"000001.SZ": df1, "600519.SH": df2},
-    factor_universe={"000001.SZ": f1, "600519.SH": f2},
-)
-# → 等权组合回测报告
-
- -
-

6. 参数优化 — OptunaEngine

-

使用预置搜索空间

-
from optimizer.engine import OptunaEngine
-from optimizer.space import rsi_revert_space, sma_cross_space
-from backtest.strategies.rsi_mean_revert import RSIMeanRevertStrategy
-
-opt_engine = OptunaEngine(engine_bt)
-
-# 优化 RSI 反转策略参数
-result = opt_engine.optimize(
-    strategy_class=RSIMeanRevertStrategy,
-    search_space=rsi_revert_space,
-    price_df=price_df,
-    factor_df=factor_df,
-    metric="sharpe",       # 优化目标: sharpe/cagr/calmar/total_return
-    n_trials=200,          # 试验次数
-)
-
- -

读取优化结果

-
print(result.summary())
-# → 最优参数: oversold=13, overbought=66
-# → 最优目标 (sharpe): 0.5985
-
-# 最优参数的回测报告
-best_report = result.best_report
-
-# 参数重要性
-for k, v in sorted(result.param_importance.items(), key=lambda x: -x[1]):
-    print(f"  {k}: {v:.4f}")
-
-# 试验记录
-trials = result.trials_df  # pd.DataFrame
-
- -

Walk-Forward 验证

-
wf_result = opt_engine.optimize_walk_forward(
-    strategy_class=RSIMeanRevertStrategy,
-    search_space=rsi_revert_space,
-    price_df=price_df,
-    factor_df=factor_df,
-    metric="sharpe",
-    n_trials=80,
-    train_window=252 * 3,   # 3 年训练
-    test_window=252,         # 1 年测试
-)
-print(wf_result.summary())
-# → 各窗口参数变化 + 整体收益
-
- -

快捷函数

-
from optimizer.presets import (
-    optimize_sma_cross,
-    optimize_rsi_revert,
-    optimize_momentum_breakout,
-)
-
-result = optimize_rsi_revert(price_df, factor_df, engine_bt, n_trials=100)
-
- -

自定义搜索空间

-
from optimizer.space import SearchSpace
-
-my_space = SearchSpace(params=[
-    {"name": "fast", "type": "int", "low": 2, "high": 30, "step": 1},
-    {"name": "slow", "type": "int", "low": 15, "high": 120, "step": 5},
-])
-result = opt_engine.optimize(MyStrategy, my_space, price_df, factor_df)
-
- -
-

7. ML 模型 — LightGBM / CatBoost

-

特征工程

-
from models.features import FeatureEngine
-
-# lookahead=5: 预测未来 5 个交易日收益
-fe = FeatureEngine(lookahead=5, label_type="regression")
-
-# 构建特征矩阵和标签
-X, y = fe.build(factor_df, price_df, fit=True)
-# → X: 标准特征矩阵(去极值 → 缺失填充 → RobustScaler)
-# → y: 未来 5 日收益率(%)
-
-print(f"特征: {X.shape[1]} 列, 样本: {X.shape[0]} 行")
-print(f"标签: mean={y.mean():.2f}%, std={y.std():.2f}%")
-
- -

数据划分

-
# 时间序列划分(前 70% 训练,后 30% 测试)
-n = len(X)
-split = int(n * 0.7)
-X_train, X_test = X.iloc[:split], X.iloc[split:]
-y_train, y_test = y.iloc[:split], y.iloc[split:]
-
-print(f"训练集: {len(X_train)} 行")
-print(f"测试集: {len(X_test)} 行")
-
- -

LightGBM 训练

-
from models.lightgbm.model import LightGBMModel
-
-model = LightGBMModel(
-    params={
-        "n_estimators": 200,
-        "learning_rate": 0.03,
-        "num_leaves": 15,
-    },
-    early_stopping=100,
-    eval_ratio=0.2,  # 20% 做验证集
-)
-
-model.fit(X_train, y_train)
-pred = model.predict(X_test)
-
-# 评估
-ic = pred.corr(y_test)
-print(f"测试集 IC: {ic:.4f}")
-
- -

CatBoost 训练

-
from models.catboost.model import CatBoostModel
-
-model = CatBoostModel(
-    params={"iterations": 200, "learning_rate": 0.03, "depth": 5},
-    eval_ratio=0.2,
-)
-model.fit(X_train, y_train)
-pred = model.predict(X_test)
-
- -

特征重要性

-
# LightGBM
-imp = model.get_feature_importance(importance_type="gain")
-print(imp.head(10))
-# → feature, importance, importance_pct
-
-# CatBoost
-imp = model.get_feature_importance()
-print(imp.head(5))
-
- -

交叉验证

-
# 5 折时间序列 CV(不 shuffle)
-cv_df = model.cv_evaluate(X_train, y_train, n_folds=5)
-print(cv_df)
-# → 各折 IC + MSE, 均值
-
- -

模型持久化

-
# 保存
-model.save("models/lightgbm_000001.pkl")
-
-# 加载
-model = LightGBMModel.load("models/lightgbm_000001.pkl")
-
- -

ML 策略回测

-
from models.backtest_integration import MLStrategy, MLBenchmark
-
-# 预测值分位 → 交易信号
-strategy = MLStrategy(
-    model=model,
-    feature_engine=fe,
-    buy_quantile=0.7,   # 预测值最高的 30% 买入
-    sell_quantile=0.3,  # 预测值最低的 30% 卖出
-    rebalance_freq=5,   # 每 5 日调仓
-)
-report = engine_bt.run(strategy, price_df, factor_df)
-
-# 多模型对比
-benchmark = MLBenchmark(
-    models=[lgb_model, cb_model],
-    feature_engine=fe,
-    price_df=test_price,
-    factor_df=test_factor,
-)
-df = benchmark.run()
-print(df)
-# → model × (IC, total_return, sharpe, win_rate, trades)
-
- -
-

8. 情绪因子 — SentimentEngine

-

配置 API Key

-

编辑 finance/.env:

-
# DashScope API(推荐)
-QWEN_API_KEY=sk-your-key-here
-QWEN_MODEL=qwen-turbo
-
-# 或本地 Ollama
-# QWEN_LOCAL_BASE_URL=http://localhost:11434/v1
-# QWEN_LOCAL_MODEL=qwen2.5:7b
-
- -

配置分析范围

-
# 按指数成分股分析(沪深300 + 中证500)
-SENTIMENT_SCOPE_TYPE=index
-SENTIMENT_SCOPE_INDEXES=000300,000905
-
-# 按板块分析
-# SENTIMENT_SCOPE_TYPE=sector
-# SENTIMENT_SCOPE_SECTORS=银行,电力设备,医药生物
-
-# 按自定义列表
-# SENTIMENT_SCOPE_TYPE=custom
-# SENTIMENT_SCOPE_CUSTOM=000001.SZ,600519.SH,300750.SZ
-
- -

使用情绪引擎

-
from factors.sentiment.sentiment_engine import SentimentEngine
-from factors.sentiment.news_source import NewsSource
-from factors.sentiment.qwen_client import QwenClient
-
-sent = SentimentEngine(dm, qwen_client=QwenClient(), news_source=NewsSource())
-
-# 单股票情绪因子
-sent_df = sent.compute("000001.SZ", max_news=20)
-# → DataFrame: (trade_date, news_sent_5, news_conf_5, sent_delta_5)
-
-# 批量计算
-results = sent.compute_batch(
-    ts_codes=["000001.SZ", "600519.SH", "300750.SZ"],
-    max_news=10,
-)
-
- -

新闻数据源

-

系统聚合三个数据源:

- - - - - - - - - - - - - - - - - - - - - - - - - -
数据源说明配置
AkShare stock_news_em东方财富个股新闻use_akshare=True
MariaDB xwlb_daily_ext新闻联播分割数据use_xwlb=True
MCP trendradar-news外部新闻聚合服务use_mcp=True
-
news = NewsSource(
-    use_akshare=True,    # 启用东方财富
-    use_xwlb=True,       # 启用新闻联播
-    use_mcp=False,       # 关闭 MCP
-)
-
-news_df = news.fetch("000001.SZ", start="20260501", end="20260603")
-# → DataFrame: date, title, content, source, url
-
- -

日期对齐机制

- -
周五新闻联播 → +1 = 周六 → align → 下周一交易日
-
- -
-

9. Agent 系统 — 命令行入口

-

注册 Agent

-
from agents.orchestrator import AgentOrchestrator
-
-engines = {
-    "dm": dm,
-    "fe": engine_fe,
-    "bt": engine_bt,
-    "opt": opt_engine,
-    "sent": sent,
-}
-
-orch = AgentOrchestrator(**engines)
-orch.setup()
-# → [Orchestrator] 已注册 4 个 Agent: ['research', 'selection', 'risk', 'report']
-
- -

CLI 命令

-
# 完整每日流程(同步行情 → 风险评估 → 选股打分 → 生成日报)
-python finance/cli/agent_cli.py daily
-
-# 今日选股 Top 15
-python finance/cli/agent_cli.py picks 15
-
-# 风险评估
-python finance/cli/agent_cli.py risk
-
-# 因子研究(IC 评估)
-python finance/cli/agent_cli.py research
-
-# 生成指定日期日报
-python finance/cli/agent_cli.py report 20260603
-
- -

每日流程输出

-
============================================================
-[Orchestrator] 每日流程 — 20260603
-============================================================
-
-[Step 1/4] 同步行情... 0 条(已是最新)
-[Step 2/4] 风险评估... high, 仓位 30%
-[Step 3/4] 股票打分... 1 只
-[Step 4/4] 生成日报... reports/daily_20260603.md
-
- -

日报输出

-

日报保存到 finance/reports/daily_YYYYMMDD.md,内容包含:

- -

编程调用

-
# 各 Agent 独立调用
-selection_result = orch.picks(date="20260603", top_n=15)
-risk_result = orch.risk_check()
-research_result = orch.run_research_cycle()
-report_result = orch.generate_report(date="20260603")
-
- -
-

10. 配置说明

-

环境变量(finance/.env)

-
# ── Qwen API ──────────────────────
-QWEN_API_KEY=sk-xxx              # DashScope API Key
-QWEN_MODEL=qwen-turbo            # 模型选择: qwen-turbo/plus/max
-
-# ── 本地 Ollama(可选) ──────────
-# QWEN_LOCAL_BASE_URL=http://localhost:11434/v1
-# QWEN_LOCAL_MODEL=qwen2.5:7b
-
-# ── 数据库 ───────────────────────
-MAC_DB_HOST=127.0.0.1
-MAC_DB_PORT=13306
-MAC_DB_USER=myquant
-MAC_DB_PASSWORD=<your-db-password>
-MAC_DB_NAME=myquant
-
-# ── 情绪分析范围 ─────────────────
-SENTIMENT_SCOPE_TYPE=index
-SENTIMENT_SCOPE_INDEXES=000300,000905
-SENTIMENT_MAX_NEWS_PER_STOCK=20
-
-# ── MCP 新闻服务(可选) ─────────
-NEWS_MCP_URL=http://192.168.1.160:3333/mcp
-
- -

回测参数

-
VectorBTEngine(
-    initial_capital=100_000,   # 初始资金(元)
-    commission=0.0003,         # 手续费(万三)
-)
-
- -

Optuna 参数

-
opt_engine.optimize(
-    n_trials=200,              # 试验次数
-    metric="sharpe",           # 优化目标
-    # 可选: cagr, calmar, total_return, return_over_dd, win_rate, profit_factor
-)
-
- -

ML 模型参数

-
# LightGBM 推荐参数
-LightGBMModel(params={
-    "n_estimators": 200,
-    "learning_rate": 0.03,
-    "num_leaves": 15,
-    "min_data_in_leaf": 20,
-    "feature_fraction": 0.7,
-    "bagging_fraction": 0.7,
-})
-
-# CatBoost 推荐参数
-CatBoostModel(params={
-    "iterations": 200,
-    "learning_rate": 0.03,
-    "depth": 5,
-    "min_data_in_leaf": 20,
-})
-
- -
-

11. 完整示例

-

示例 1:快速回测

-
import sys; sys.path.insert(0, "finance")
-
-from data.data_manager import DataManager
-from factors.registry import get_factor
-from factors.engine import FactorEngine
-from backtest.vectorbt.engine import VectorBTEngine
-from backtest.strategies.rsi_mean_revert import RSIMeanRevertStrategy
-
-# 数据
-dm = DataManager(); dm.init_db()
-price = dm.get_daily("000001.SZ").set_index("trade_date")
-
-# 因子
-engine_fe = FactorEngine(dm)
-factor_df = engine_fe.compute("000001.SZ", [get_factor("rsi_14")])
-
-# 回测
-engine_bt = VectorBTEngine()
-report = engine_bt.run(
-    RSIMeanRevertStrategy(oversold=30, overbought=70),
-    price, factor_df,
-)
-print(report.summary())
-
- -

示例 2:策略寻优 + Walk-Forward

-
from optimizer.engine import OptunaEngine
-from optimizer.space import rsi_revert_space
-
-opt_engine = OptunaEngine(engine_bt)
-
-# 寻优
-result = opt_engine.optimize(
-    RSIMeanRevertStrategy, rsi_revert_space,
-    price, factor_df, metric="sharpe", n_trials=200,
-)
-print(result.summary())
-
-# Walk-Forward 验证
-wf = opt_engine.optimize_walk_forward(
-    RSIMeanRevertStrategy, rsi_revert_space,
-    price, factor_df, n_trials=80,
-    train_window=756, test_window=252,
-)
-print(wf.summary())
-
- -

示例 3:ML 训练 + 回测

-
from models.features import FeatureEngine
-from models.lightgbm.model import LightGBMModel
-from models.backtest_integration import MLStrategy
-
-# 特征工程
-fe = FeatureEngine(lookahead=5)
-X, y = fe.build(factor_df, price, fit=True)
-split = int(len(X) * 0.7)
-
-# 训练
-model = LightGBMModel(params={"n_estimators": 200, "learning_rate": 0.03})
-model.fit(X.iloc[:split], y.iloc[:split])
-
-# 回测
-strategy = MLStrategy(model, fe)
-report = engine_bt.run(strategy, price, factor_df)
-print(report.summary())
-print(model.get_feature_importance().head(5))
-
- -

示例 4:每日 Agent 运行

-
from agents.orchestrator import AgentOrchestrator
-
-orch = AgentOrchestrator(
-    dm=dm, fe=engine_fe, bt=engine_bt, opt=opt_engine, sent=sent,
-)
-orch.setup()
-results = orch.run_daily()
-
-# 获取结果
-sel = results["selection"]
-risk = results["risk"]
-report_path = results["report"]["report_path"]
-print(f"日报: {report_path}")
-
- -
-

12. 常见问题

-

Q: SSH 隧道连接失败?

-
# 检查端口
-lsof -i :13306 | grep LISTEN
-
-# 重新建立
-bash shared/script/autossh.sh
-
- -

Q: AkShare 返回 RemoteDisconnected?

-

这是 AkShare 的 curl_cffi 在连续请求时偶发的连接问题。系统已内置 3 次递增间隔重试 + fallback 机制,通常第 2-3 次重试会成功。如果持续失败:

- -

Q: 因子计算结果全是 NaN?

- -

Q: 回测结果为 0 笔交易?

- -

Q: 模型训练只有 2 棵树?

-

当验证集损失不下降时,早停会在很少的迭代后触发。这是单股票预测的正常现象(信号噪声比低)。建议:

- -

Q: 情绪因子返回空?

- -

Q: 日报中选股为空?

-

日报只对 DB 中有日线缓存的股票打分。需要先同步目标股票池的数据:

-
# 同步单只
-dm.sync_daily("000001.SZ")
-
-# 按范围批量同步(需先配置 SENTIMENT_SCOPE)
-codes = sent.get_scope_stocks()
-for code in codes[:10]:
-    dm.sync_daily(code)
-
- -
-

13. CLI 脚本参考

-

所有脚本位于 finance/cli/,需在项目根目录或 finance/ 下运行。

-
-

13.1 Agent 系统入口 — agent_cli.py

-
cd finance && python cli/agent_cli.py <命令> [参数]
-
- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -
命令说明示例
daily [DATE]完整每日流程(增量同步已缓存→评估风险→选股→日报)agent_cli.py daily
picks [N] [DATE]多因子选股 Top N(需已缓存)agent_cli.py picks 15
risk市场风险评估(等级、仓位、止损)agent_cli.py risk
research因子发现:遍历因子计算 IC/IC_IR 排名agent_cli.py research
report [DATE]生成日报(含三指数行情+选股+情绪+风险评估)agent_cli.py report
warmup [N]首次批量预热范围股票到 DB 缓存(每批 N 只,默认 50)agent_cli.py warmup 50
-

daily 流程:

-
[Step 1/4] 增量同步 → 只更新已缓存股票(最新则 0.04s 跳过)
-              → 未缓存提示:运行 'agent_cli.py warmup' 首次预热
-[Step 2/4] 风险评估 → high/medium/low + 仓位建议 + 预警
-[Step 3/4] 股票打分 → DB 缓存命中率 + 多因子等权打分 → Top 15
-[Step 4/4] 日报生成 → 三指数行情 (Tushare) + 情绪摘要 + 风险预警
-                       → reports/daily_YYYYMMDD.md
-
- -

数据源优先级:Tushare → AkShare(.env 配置 TUSHARE_TOKEN)

-
-

13.2 数据层验证 — demo_data_manager.py

-
python cli/demo_data_manager.py [--ts_code CODE] [--start YYYYMMDD]
-
- - - - - - - - - - - - - - - - - - - - - -
参数默认值说明
--ts_code000001.SZ测试股票代码
--start20250101起始日期 YYYYMMDD
-

5 步验证:数据库连接 → 建表 → 股票列表 → 日线获取(双源fallback) → 增量同步。

-
-

13.3 因子引擎验证 — demo_factor_engine.py

-
python cli/demo_factor_engine.py [--ts_code CODE] [--ts_code2 CODE]
-
- - - - - - - - - - - - - - - - - - - - - -
参数默认值说明
--ts_code000001.SZ测试股票代码
--ts_code2600519.SH截面测试第二只股票
-

验证:因子注册表(12分类/34因子)→ 技术因子计算(describe统计) → 基本面因子(ROE/PE/PB/EP) → NaN 覆盖率检查 → 双股票截面因子。

-
-

13.4 回测引擎验证 — demo_backtest.py

-
python cli/demo_backtest.py [--ts_code CODE]
-
- - - - - - - - - - - - - - - - -
参数默认值说明
--ts_code000001.SZ回测股票代码
-

测试 5 个内置策略:

- - - - - - - - - - - - - - - - - - - - - - - - - - - - - -
策略参数
SMACrossStrategy(5,20) / (10,60)
RSIMeanRevertStrategy(30,70) / (20,80)
MomentumBreakoutStrategylookback=20
FactorCrossStrategymomentum_20 > 0
FactorRotationStrategymomentum top 20%
-
-

13.5 参数优化验证 — demo_optimizer.py

-
python cli/demo_optimizer.py [--ts_code CODE] [--trials N]
-
- - - - - - - - - - - - - - - - - - - - - -
参数默认值说明
--ts_code000001.SZ回测股票代码
--trials200Optuna 试验次数
-

对 RSI 反转策略执行参数寻优 + Walk-Forward 验证。输出最优 vs 默认对比表 + 参数重要性排序。

-
-

13.6 ML 模型验证 — demo_ml.py

-
python cli/demo_ml.py [--ts_code CODE] [--lookahead N]
-
- - - - - - - - - - - - - - - - - - - - - -
参数默认值说明
--ts_code000001.SZ训练股票代码
--lookahead5预测未来 N 日收益
-

完整 ML pipeline:特征工程(25因子→Winsorize→RobustScaler) → LightGBM训练(IC/CV) → CatBoost训练 → MLBenchmark对比(IC/收益/夏普/胜率)。

-
-

13.7 情绪因子快速验证 — demo_sentiment.py

-
python cli/demo_sentiment.py [--ts_code CODE] [--no-qwen]
-
- - - - - - - - - - - - - - - - - - - - - -
参数默认值说明
--ts_code000001.SZ测试股票代码
--no-qwenflag跳过 Qwen API 调用
-

6 步验证:新闻数据源(三源聚合) → 日期对齐 → Qwen 客户端状态 → SentimentEngine全链路 → 分析范围解析。

-

适合快速检查情绪因子系统是否就绪。

-
-

13.8 情绪因子详细演示 — demo_sentiment_detail.py

-
python cli/demo_sentiment_detail.py [选项]
-
- -

最详细的情绪因子脚本,支持完整命令行参数和逐步输出。

- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -
参数类型默认值说明
--ts_codestr000001.SZ股票代码,多个用逗号分隔
--datestr今天目标日期 YYYYMMDD
--startstrdate-30天起始日期 YYYYMMDD
--endstrdate结束日期 YYYYMMDD
--scope-typestr-分析范围:index/sector/custom/all
--scope-indexesstr000300指数代码(逗号分隔)
--scope-sectorsstr-板块名称(逗号分隔)
--max-newsint.env 配置最大新闻条数
--max-analyzeint50Qwen API 分析最大条数(控制成本)
--no-xwlbflag-禁用新闻联播数据源
--no-akshareflag-禁用东方财富数据源
--no-mcpflag-禁用 MCP 数据源
--sourcestr-仅用指定数据源:xwlb/akshare/mcp
--no-qwenflag-跳过 Qwen API 调用(仅演示数据流)
-

使用示例:

-
# 默认演示(000001.SZ,最近30天,全数据源)
-python cli/demo_sentiment_detail.py
-
-# 指定股票和日期
-python cli/demo_sentiment_detail.py --ts_code 600519.SH --date 20260603
-
-# 多股票 + 日期范围
-python cli/demo_sentiment_detail.py --ts_code 000001.SZ,300316.SZ --start 20260501 --end 20260603
-
-# 按指数成分股分析
-python cli/demo_sentiment_detail.py --scope-type index --scope-indexes 000300
-
-# 按板块分析
-python cli/demo_sentiment_detail.py --scope-type sector --scope-sectors 银行,电力设备
-
-# 只看东方财富新闻,不调用 Qwen
-python cli/demo_sentiment_detail.py --source akshare --no-qwen --max-news 20
-
- -

输出 6 步详情:

-
Step 0: 初始化引擎(显示数据源、API状态、范围、Tushare可用性)
-Step 1: 按数据源分别拉取新闻(xwlb/AkShare/MCP 各自数量 + 双源fallback)
-Step 2: 新闻详情(按来源分开展示标题/内容/链接)
-Step 3: 日期对齐(xwlb +1day偏移 + 非交易日对齐 + DB缓存检查→sync补齐)
-Step 4: Qwen 情绪分析(每条新闻的分数/置信度/主题/来源标签)
-Step 5: 因子计算(weighted sent / confidence-weighted / momentum + 公式说明)
-Step 6: 结果输出(因子值表 + 历史统计 + 每新闻情绪贡献明细)
-
- -
-

13.9 脚本一览

- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -
脚本参数用途数据源 fallback
agent_cli.py子命令 + 参数日常操作入口✅
demo_data_manager.py--ts_code --startSprint 0 验证✅
demo_factor_engine.py--ts_code --ts_code2Sprint 1 验证✅
demo_backtest.py--ts_codeSprint 2 验证✅
demo_optimizer.py--ts_code --trialsSprint 3 验证✅
demo_ml.py--ts_code --lookaheadSprint 4 验证✅
demo_sentiment.py--ts_code --no-qwenSprint 5 快速验证✅
demo_sentiment_detail.py14 个 argparse 参数Sprint 5 详细演示✅
- - - \ No newline at end of file diff --git a/docs/usage.md b/docs/usage.md index 888d021..32c3285 100644 --- a/docs/usage.md +++ b/docs/usage.md @@ -1,6 +1,6 @@ # cc-cursor 使用指南 -Mac Mini 单机量化研究平台,覆盖数据获取 → 因子计算 → 回测 → 参数优化 → ML 模型 → 情绪因子 → Agent 系统 全链路。 +Mac Mini 单机量化研究平台,覆盖数据获取 → 因子计算 → 回测 → 参数优化 → ML 模型 → 情绪因子 → Agent 系统全链路。 --- @@ -32,11 +32,8 @@ Mac Mini 单机量化研究平台,覆盖数据获取 → 因子计算 → 回 ### Python 环境 ```bash -# 激活 conda 环境 -conda activate quant - -# 确认 Python 版本 -python --version # → 3.11.13 +conda activate quant # Python 3.11.13 +python --version # → 3.11.13 ``` ### 核心依赖 @@ -56,7 +53,7 @@ python --version # → 3.11.13 ### 项目路径设置 -所有代码从 `finance/` 目录运行。Python 脚本开头加入: +从 `finance/` 目录运行代码。脚本开头加入: ```python import sys @@ -69,8 +66,6 @@ sys.path.insert(0, "/path/to/cc-cursor/finance") ### 建立 SSH 隧道 -系统通过 SSH 隧道连接远程 MariaDB: - ```bash bash shared/script/autossh.sh ``` @@ -79,7 +74,6 @@ bash shared/script/autossh.sh ```bash lsof -i :13306 | grep LISTEN -# → ssh ... localhost:13306 (LISTEN) ... ``` ### 连接信息 @@ -94,24 +88,14 @@ Database: myquant ### 数据库表 -所有表使用 `mac_` 前缀,与已有表隔离: +所有表使用 `mac_` 前缀: | 表名 | 内容 | 说明 | |------|------|------| | `mac_stock_basic` | A 股列表 | 5,524 只股票 | | `mac_stock_daily` | 日线行情 | 按需同步 | | `mac_stock_financial` | 财务指标 | 同花顺核心指标 | - -### 测试连接 - -```python -from database.connection import test_connection - -if test_connection(): - print("数据库连接成功") -else: - print("请先建立 SSH 隧道: bash shared/script/autossh.sh") -``` +| `mac_report` | 报告持久化 | 日报归档 | --- @@ -129,48 +113,32 @@ dm.init_db() # 首次使用创建表(幂等操作) ### 获取股票列表 ```python -# 从 DB 缓存读取(已缓存的 5,524 只 A 股) -stocks = dm.get_stock_list() -# → DataFrame: index=ts_code, columns=[name, area, industry, ...] - -# 强制从 AkShare 刷新 -stocks = dm.get_stock_list(force_refresh=True) +stocks = dm.get_stock_list() # → 5,524 只 A 股 +stocks = dm.get_stock_list(force_refresh=True) # 强制刷新 ``` ### 获取日线数据 ```python -# 获取单只股票日线(DB 缓存优先,缺失自动补拉) -daily = dm.get_daily("000001.SZ") -# → DataFrame: trade_date, open, high, low, close, vol, amount, ... - -# 指定日期范围 +daily = dm.get_daily("000001.SZ") # 全量日线 daily = dm.get_daily("000001.SZ", start="20240101", end="20241231") - -# 直接用索引 -price = daily.set_index("trade_date").sort_index() -close = price["close"] ``` ### 获取财务数据 ```python fina = dm.get_financial("000001.SZ") -# → DataFrame: end_date, eps, bvps, roe, net_profit_margin, debt_to_assets, ... -# 数据源: stock_financial_abstract_ths(同花顺) -# 覆盖: 主板/创业板/科创板 +# → DataFrame: end_date, eps, bvps, roe, net_profit_margin, debt_to_assets ``` ### 数据同步 ```python -# 增量同步:从 DB 最新日期到今天的缺失数据 -n = dm.sync_daily("000001.SZ") - -# 批量同步全部股票(谨慎使用,耗时长) -total = dm.sync_all_daily() +n = dm.sync_daily("000001.SZ") # 增量同步到最新 ``` +> 数据源优先级:Tushare(优先)→ AkShare(fallback)。DB 缓存优先。 + --- ## 4. 因子引擎 — FactorEngine @@ -180,32 +148,18 @@ total = dm.sync_all_daily() ```python from factors.registry import get_factor, list_factors, list_categories -# 查看所有因子分类 -print(list_categories()) -# → ['动量', 'RSI', 'MACD', '量价', '布林', 'ATR', '均线', '波动率', '换手率', '振幅', '基本面', '情绪'] - -# 查看某个分类下的因子 -print(list_factors("RSI")) -# → ['rsi_7', 'rsi_14'] - -# 查看全部因子 -all_factors = list_factors() -print(len(all_factors)) -# → 34 +print(list_categories()) # → 12 个分类 +print(list_factors("RSI")) # → ['rsi_7', 'rsi_14'] +print(len(list_factors())) # → 34 个因子 ``` ### 创建因子实例 ```python -# 按名称获取(使用默认参数) -factor = get_factor("momentum_20") # 20 日动量 -factor = get_factor("rsi_14") # 14 日 RSI -factor = get_factor("roe") # ROE 基本面因子 -factor = get_factor("news_sent_5") # 5 日新闻情绪因子 - -# 自定义参数 -from factors.technical.momentum import MomentumFactor -factor = MomentumFactor(period=60) +factor = get_factor("momentum_20") # 20 日动量 +factor = get_factor("rsi_14") # 14 日 RSI +factor = get_factor("roe") # ROE 基本面因子 +factor = get_factor("news_sent_5") # 5 日新闻情绪因子 ``` ### 计算因子 @@ -214,64 +168,37 @@ factor = MomentumFactor(period=60) from factors.engine import FactorEngine engine_fe = FactorEngine(dm) - -# 单股票多因子 factors = [ get_factor("momentum_20"), get_factor("rsi_14"), get_factor("volatility_20"), - get_factor("ma_dev_20"), ] factor_df = engine_fe.compute("000001.SZ", factors) -# → DataFrame: index=trade_date, columns=[momentum_20, rsi_14, volatility_20, ma_dev_20] - -# 查看因子值 -print(factor_df.tail()) -print(factor_df.describe()) +# → DataFrame: index=trade_date, columns=[momentum_20, rsi_14, volatility_20] ``` ### 截面因子 ```python -# 计算多只股票在某一天的因子值 cross = engine_fe.compute_universe( factors=[get_factor("momentum_20"), get_factor("rsi_14")], date="20250630", ts_codes=["000001.SZ", "600519.SH", "300750.SZ"], ) -# → DataFrame: index=ts_code, columns=[momentum_20, rsi_14] -``` - -### 因子质量检查 - -```python -# 查看 NaN 率 -total = len(factor_df) -for col in factor_df.columns: - nan_pct = factor_df[col].isna().sum() / total * 100 - print(f"{col}: NaN {nan_pct:.1f}%") -# 正常范围: 技术因子 0.3%-2.1%, 基本面因子 0% ``` --- ## 5. 回测引擎 — VectorBTEngine -### 参数 - -- 初始资金:100,000 元 -- 手续费:0.03%(万三) -- 方向:只做多 +- 初始资金:100,000 元 | 手续费:0.03%(万三) | 方向:只做多 ### 创建引擎 ```python from backtest.vectorbt.engine import VectorBTEngine -engine_bt = VectorBTEngine( - initial_capital=100_000, - commission=0.0003, -) +engine_bt = VectorBTEngine(initial_capital=100_000, commission=0.0003) ``` ### 使用内置策略 @@ -279,30 +206,10 @@ engine_bt = VectorBTEngine( ```python from backtest.strategies.rsi_mean_revert import RSIMeanRevertStrategy -# 创建策略 strategy = RSIMeanRevertStrategy(oversold=30, overbought=70) - -# 运行回测 report = engine_bt.run(strategy, price_df, factor_df) -``` - -### 读取回测报告 - -```python -# 一行摘要 print(report.summary()) -# → 收益=29.4% 年化=4.3% 回撤=-19.1% 夏普=0.37 胜率=77.1% 交易=70笔 - -# 字典格式 -metrics = report.to_dict() -# → {'total_return': 29.4, 'cagr': 4.3, 'sharpe_ratio': 0.37, ...} - -# 获取净值曲线 -equity = report.equity_curve # pd.Series -drawdown = report.drawdown_curve # pd.Series - -# 逐笔交易 -trades = report.trades_df # pd.DataFrame +# → 收益=29.4% 年化=4.3% 回撤=-19.1% 夏普=0.37 胜率=77.1% ``` ### 内置策略清单 @@ -328,26 +235,15 @@ class MyStrategy(BaseStrategy): self.param_a = param_a def generate_signals(self, factor_df): - # factor_df 包含因子值和 close 列 - # 返回: 1=买入, 0=卖出, -1=持有 signals = pd.Series(-1, index=factor_df.index) - signals[factor_df["rsi_14"] < 30] = 1 # RSI 超卖买入 - signals[factor_df["rsi_14"] > 70] = 0 # RSI 超买卖出 + signals[factor_df["rsi_14"] < 30] = 1 # 超卖买入 + signals[factor_df["rsi_14"] > 70] = 0 # 超买卖出 return signals - -report = engine_bt.run(MyStrategy(param_a=20), price_df, factor_df) ``` -### 截面回测(多股票) +### 回测报告字段 -```python -report_xs = engine_bt.run_cross_section( - strategy, - price_universe={"000001.SZ": df1, "600519.SH": df2}, - factor_universe={"000001.SZ": f1, "600519.SH": f2}, -) -# → 等权组合回测报告 -``` +`total_return`, `cagr`, `max_drawdown`, `sharpe_ratio`, `calmar_ratio`, `annual_volatility`, `win_rate`, `profit_factor`, `total_trades`, `avg_hold_days`, `equity_curve`, `drawdown_curve`, `monthly_returns`, `trades_df` --- @@ -357,38 +253,21 @@ report_xs = engine_bt.run_cross_section( ```python from optimizer.engine import OptunaEngine -from optimizer.space import rsi_revert_space, sma_cross_space -from backtest.strategies.rsi_mean_revert import RSIMeanRevertStrategy +from optimizer.space import rsi_revert_space opt_engine = OptunaEngine(engine_bt) -# 优化 RSI 反转策略参数 result = opt_engine.optimize( strategy_class=RSIMeanRevertStrategy, search_space=rsi_revert_space, price_df=price_df, factor_df=factor_df, - metric="sharpe", # 优化目标: sharpe/cagr/calmar/total_return - n_trials=200, # 试验次数 + metric="sharpe", # sharpe/cagr/calmar/total_return + n_trials=200, ) -``` - -### 读取优化结果 - -```python print(result.summary()) # → 最优参数: oversold=13, overbought=66 # → 最优目标 (sharpe): 0.5985 - -# 最优参数的回测报告 -best_report = result.best_report - -# 参数重要性 -for k, v in sorted(result.param_importance.items(), key=lambda x: -x[1]): - print(f" {k}: {v:.4f}") - -# 试验记录 -trials = result.trials_df # pd.DataFrame ``` ### Walk-Forward 验证 @@ -397,27 +276,10 @@ trials = result.trials_df # pd.DataFrame wf_result = opt_engine.optimize_walk_forward( strategy_class=RSIMeanRevertStrategy, search_space=rsi_revert_space, - price_df=price_df, - factor_df=factor_df, - metric="sharpe", - n_trials=80, - train_window=252 * 3, # 3 年训练 - test_window=252, # 1 年测试 + price_df=price_df, factor_df=factor_df, + metric="sharpe", n_trials=80, + train_window=756, test_window=252, ) -print(wf_result.summary()) -# → 各窗口参数变化 + 整体收益 -``` - -### 快捷函数 - -```python -from optimizer.presets import ( - optimize_sma_cross, - optimize_rsi_revert, - optimize_momentum_breakout, -) - -result = optimize_rsi_revert(price_df, factor_df, engine_bt, n_trials=100) ``` ### 自定义搜索空间 @@ -429,7 +291,6 @@ my_space = SearchSpace(params=[ {"name": "fast", "type": "int", "low": 2, "high": 30, "step": 1}, {"name": "slow", "type": "int", "low": 15, "high": 120, "step": 5}, ]) -result = opt_engine.optimize(MyStrategy, my_space, price_df, factor_df) ``` --- @@ -441,29 +302,9 @@ result = opt_engine.optimize(MyStrategy, my_space, price_df, factor_df) ```python from models.features import FeatureEngine -# lookahead=5: 预测未来 5 个交易日收益 fe = FeatureEngine(lookahead=5, label_type="regression") - -# 构建特征矩阵和标签 X, y = fe.build(factor_df, price_df, fit=True) -# → X: 标准特征矩阵(去极值 → 缺失填充 → RobustScaler) -# → y: 未来 5 日收益率(%) - -print(f"特征: {X.shape[1]} 列, 样本: {X.shape[0]} 行") -print(f"标签: mean={y.mean():.2f}%, std={y.std():.2f}%") -``` - -### 数据划分 - -```python -# 时间序列划分(前 70% 训练,后 30% 测试) -n = len(X) -split = int(n * 0.7) -X_train, X_test = X.iloc[:split], X.iloc[split:] -y_train, y_test = y.iloc[:split], y.iloc[split:] - -print(f"训练集: {len(X_train)} 行") -print(f"测试集: {len(X_test)} 行") +# → Winsorize(1%/99%) → ffill → median fill → RobustScaler → 标签计算 ``` ### LightGBM 训练 @@ -472,21 +313,12 @@ print(f"测试集: {len(X_test)} 行") from models.lightgbm.model import LightGBMModel model = LightGBMModel( - params={ - "n_estimators": 200, - "learning_rate": 0.03, - "num_leaves": 15, - }, - early_stopping=100, - eval_ratio=0.2, # 20% 做验证集 + params={"n_estimators": 200, "learning_rate": 0.03, "num_leaves": 15}, + eval_ratio=0.2, ) - model.fit(X_train, y_train) pred = model.predict(X_test) - -# 评估 ic = pred.corr(y_test) -print(f"测试集 IC: {ic:.4f}") ``` ### CatBoost 训练 @@ -499,39 +331,6 @@ model = CatBoostModel( eval_ratio=0.2, ) model.fit(X_train, y_train) -pred = model.predict(X_test) -``` - -### 特征重要性 - -```python -# LightGBM -imp = model.get_feature_importance(importance_type="gain") -print(imp.head(10)) -# → feature, importance, importance_pct - -# CatBoost -imp = model.get_feature_importance() -print(imp.head(5)) -``` - -### 交叉验证 - -```python -# 5 折时间序列 CV(不 shuffle) -cv_df = model.cv_evaluate(X_train, y_train, n_folds=5) -print(cv_df) -# → 各折 IC + MSE, 均值 -``` - -### 模型持久化 - -```python -# 保存 -model.save("models/lightgbm_000001.pkl") - -# 加载 -model = LightGBMModel.load("models/lightgbm_000001.pkl") ``` ### ML 策略回测 @@ -539,28 +338,20 @@ model = LightGBMModel.load("models/lightgbm_000001.pkl") ```python from models.backtest_integration import MLStrategy, MLBenchmark -# 预测值分位 → 交易信号 -strategy = MLStrategy( - model=model, - feature_engine=fe, - buy_quantile=0.7, # 预测值最高的 30% 买入 - sell_quantile=0.3, # 预测值最低的 30% 卖出 - rebalance_freq=5, # 每 5 日调仓 -) +strategy = MLStrategy(model, fe, buy_quantile=0.7, sell_quantile=0.3, rebalance_freq=5) report = engine_bt.run(strategy, price_df, factor_df) # 多模型对比 -benchmark = MLBenchmark( - models=[lgb_model, cb_model], - feature_engine=fe, - price_df=test_price, - factor_df=test_factor, -) -df = benchmark.run() -print(df) -# → model × (IC, total_return, sharpe, win_rate, trades) +benchmark = MLBenchmark([lgb_model, cb_model], fe, price_df, factor_df) +df = benchmark.run() # → model × (IC, return, sharpe, win_rate, trades) ``` +### 重要约束 + +- lookahead 固定,不输入模型(防目标泄露) +- 交叉验证用 TimeSeriesSplit(不 shuffle) +- 特征工程严禁使用未来数据 + --- ## 8. 情绪因子 — SentimentEngine @@ -570,151 +361,77 @@ print(df) 编辑 `finance/.env`: ```bash -# DashScope API(推荐) QWEN_API_KEY=sk-your-key-here QWEN_MODEL=qwen-turbo - -# 或本地 Ollama -# QWEN_LOCAL_BASE_URL=http://localhost:11434/v1 -# QWEN_LOCAL_MODEL=qwen2.5:7b -``` - -### 配置分析范围 - -```bash -# 按指数成分股分析(沪深300 + 中证500) -SENTIMENT_SCOPE_TYPE=index -SENTIMENT_SCOPE_INDEXES=000300,000905 - -# 按板块分析 -# SENTIMENT_SCOPE_TYPE=sector -# SENTIMENT_SCOPE_SECTORS=银行,电力设备,医药生物 - -# 按自定义列表 -# SENTIMENT_SCOPE_TYPE=custom -# SENTIMENT_SCOPE_CUSTOM=000001.SZ,600519.SH,300750.SZ ``` ### 使用情绪引擎 ```python from factors.sentiment.sentiment_engine import SentimentEngine -from factors.sentiment.news_source import NewsSource -from factors.sentiment.qwen_client import QwenClient -sent = SentimentEngine(dm, qwen_client=QwenClient(), news_source=NewsSource()) - -# 单股票情绪因子 +sent = SentimentEngine(dm) sent_df = sent.compute("000001.SZ", max_news=20) # → DataFrame: (trade_date, news_sent_5, news_conf_5, sent_delta_5) - -# 批量计算 -results = sent.compute_batch( - ts_codes=["000001.SZ", "600519.SH", "300750.SZ"], - max_news=10, -) ``` ### 新闻数据源 -系统聚合三个数据源: - -| 数据源 | 说明 | 配置 | -|--------|------|------| -| AkShare `stock_news_em` | 东方财富个股新闻 | `use_akshare=True` | -| MariaDB `xwlb_daily_ext` | 新闻联播分割数据 | `use_xwlb=True` | -| MCP `trendradar-news` | 外部新闻聚合服务 | `use_mcp=True` | - -```python -news = NewsSource( - use_akshare=True, # 启用东方财富 - use_xwlb=True, # 启用新闻联播 - use_mcp=False, # 关闭 MCP -) - -news_df = news.fetch("000001.SZ", start="20260501", end="20260603") -# → DataFrame: date, title, content, source, url -``` +| 数据源 | 说明 | +|--------|------| +| AkShare `stock_news_em` | 东方财富个股新闻 | +| MariaDB `xwlb_daily_ext` | 新闻联播分割数据 | +| MCP `trendradar-news` | 外部新闻聚合服务 | ### 日期对齐机制 -- **AkShare 新闻**:`发布时间` 直接保留 → `align_news_to_trading_days` 对齐到最近交易日 -- **新闻联播**:`news_date + 1 day`(晚间播出 → 次日市场影响)→ 对齐到交易日 - -``` -周五新闻联播 → +1 = 周六 → align → 下周一交易日 -``` +- 新闻联播:`news_date + 1 day`(晚间播出 → 次日市场影响) +- 非交易日 → 对齐到最近交易日 --- ## 9. Agent 系统 — 命令行入口 -### 注册 Agent - -```python -from agents.orchestrator import AgentOrchestrator - -engines = { - "dm": dm, - "fe": engine_fe, - "bt": engine_bt, - "opt": opt_engine, - "sent": sent, -} - -orch = AgentOrchestrator(**engines) -orch.setup() -# → [Orchestrator] 已注册 4 个 Agent: ['research', 'selection', 'risk', 'report'] -``` - ### CLI 命令 ```bash -# 完整每日流程(同步行情 → 风险评估 → 选股打分 → 生成日报) -python finance/cli/agent_cli.py daily - -# 今日选股 Top 15 -python finance/cli/agent_cli.py picks 15 - -# 风险评估 -python finance/cli/agent_cli.py risk - -# 因子研究(IC 评估) -python finance/cli/agent_cli.py research - -# 生成指定日期日报 -python finance/cli/agent_cli.py report 20260603 +python finance/cli/agent_cli.py daily # 5 步完整流程 +python finance/cli/agent_cli.py picks 15 # 选股 Top 15 +python finance/cli/agent_cli.py risk # 风险评估 +python finance/cli/agent_cli.py research # 因子研究(IC 评估) +python finance/cli/agent_cli.py report 20260603 # 生成日报 +python finance/cli/agent_cli.py warmup 50 # 首次预热缓存 ``` -### 每日流程输出 +### 每日流程 ``` -============================================================ -[Orchestrator] 每日流程 — 20260603 -============================================================ - -[Step 1/4] 同步行情... 0 条(已是最新) -[Step 2/4] 风险评估... high, 仓位 30% -[Step 3/4] 股票打分... 1 只 -[Step 4/4] 生成日报... reports/daily_20260603.md +[Step 1/5] 增量同步 → 只更新已缓存股票 +[Step 2/5] 风险评估 → RiskAgent: high/medium/low + 仓位 +[Step 3/5] 股票打分 → SelectionAgent: 多因子等权打分 +[Step 4/5] 情绪因子 → SentimentEngine.compute() +[Step 5/5] 生成日报 → ReportAgent: .md + .html + 解读 ``` -### 日报输出 +### 4 个 Agent -日报保存到 `finance/reports/daily_YYYYMMDD.md`,内容包含: +| Agent | 职责 | +|-------|------| +| ResearchAgent | 因子 IC/IC_IR 评估 | +| SelectionAgent | 多因子股票打分(等权) | +| RiskAgent | 波动率+回撤→仓位建议 | +| ReportAgent | 市场+选股+情绪+风险→日报 | -- **市场概览**:上证/深证/创业板 收盘价、涨跌幅、5日/20日趋势 -- **今日推荐**:TOP 15 股票打分排名 -- **风险评估**:风险等级、建议仓位、止损线、预警 +### Demo 验证脚本 -### 编程调用 - -```python -# 各 Agent 独立调用 -selection_result = orch.picks(date="20260603", top_n=15) -risk_result = orch.risk_check() -research_result = orch.run_research_cycle() -report_result = orch.generate_report(date="20260603") +```bash +python finance/cli/demo_data_manager.py --ts_code 600519.SH +python finance/cli/demo_factor_engine.py --ts_code 300750.SZ +python finance/cli/demo_backtest.py --ts_code 000001.SZ +python finance/cli/demo_optimizer.py --ts_code 000001.SZ --trials 100 +python finance/cli/demo_ml.py --ts_code 000001.SZ --lookahead 5 +python finance/cli/demo_sentiment.py --ts_code 600519.SH +python finance/cli/demo_sentiment_detail.py --ts_code 600519.SH --date 20260603 ``` --- @@ -726,11 +443,7 @@ report_result = orch.generate_report(date="20260603") ```bash # ── Qwen API ────────────────────── QWEN_API_KEY=sk-xxx # DashScope API Key -QWEN_MODEL=qwen-turbo # 模型选择: qwen-turbo/plus/max - -# ── 本地 Ollama(可选) ────────── -# QWEN_LOCAL_BASE_URL=http://localhost:11434/v1 -# QWEN_LOCAL_MODEL=qwen2.5:7b +QWEN_MODEL=qwen-turbo # qwen-turbo/plus/max # ── 数据库 ─────────────────────── MAC_DB_HOST=127.0.0.1 @@ -739,54 +452,13 @@ MAC_DB_USER=myquant MAC_DB_PASSWORD= MAC_DB_NAME=myquant +# ── Tushare ────────────────────── +TUSHARE_TOKEN=your_token_here + # ── 情绪分析范围 ───────────────── SENTIMENT_SCOPE_TYPE=index SENTIMENT_SCOPE_INDEXES=000300,000905 SENTIMENT_MAX_NEWS_PER_STOCK=20 - -# ── MCP 新闻服务(可选) ───────── -NEWS_MCP_URL=http://192.168.1.160:3333/mcp -``` - -### 回测参数 - -```python -VectorBTEngine( - initial_capital=100_000, # 初始资金(元) - commission=0.0003, # 手续费(万三) -) -``` - -### Optuna 参数 - -```python -opt_engine.optimize( - n_trials=200, # 试验次数 - metric="sharpe", # 优化目标 - # 可选: cagr, calmar, total_return, return_over_dd, win_rate, profit_factor -) -``` - -### ML 模型参数 - -```python -# LightGBM 推荐参数 -LightGBMModel(params={ - "n_estimators": 200, - "learning_rate": 0.03, - "num_leaves": 15, - "min_data_in_leaf": 20, - "feature_fraction": 0.7, - "bagging_fraction": 0.7, -}) - -# CatBoost 推荐参数 -CatBoostModel(params={ - "iterations": 200, - "learning_rate": 0.03, - "depth": 5, - "min_data_in_leaf": 20, -}) ``` --- @@ -804,20 +476,14 @@ from factors.engine import FactorEngine from backtest.vectorbt.engine import VectorBTEngine from backtest.strategies.rsi_mean_revert import RSIMeanRevertStrategy -# 数据 dm = DataManager(); dm.init_db() price = dm.get_daily("000001.SZ").set_index("trade_date") -# 因子 engine_fe = FactorEngine(dm) factor_df = engine_fe.compute("000001.SZ", [get_factor("rsi_14")]) -# 回测 engine_bt = VectorBTEngine() -report = engine_bt.run( - RSIMeanRevertStrategy(oversold=30, overbought=70), - price, factor_df, -) +report = engine_bt.run(RSIMeanRevertStrategy(30, 70), price, factor_df) print(report.summary()) ``` @@ -828,21 +494,17 @@ from optimizer.engine import OptunaEngine from optimizer.space import rsi_revert_space opt_engine = OptunaEngine(engine_bt) - -# 寻优 result = opt_engine.optimize( RSIMeanRevertStrategy, rsi_revert_space, price, factor_df, metric="sharpe", n_trials=200, ) print(result.summary()) -# Walk-Forward 验证 wf = opt_engine.optimize_walk_forward( RSIMeanRevertStrategy, rsi_revert_space, price, factor_df, n_trials=80, train_window=756, test_window=252, ) -print(wf.summary()) ``` ### 示例 3:ML 训练 + 回测 @@ -852,38 +514,15 @@ from models.features import FeatureEngine from models.lightgbm.model import LightGBMModel from models.backtest_integration import MLStrategy -# 特征工程 fe = FeatureEngine(lookahead=5) X, y = fe.build(factor_df, price, fit=True) split = int(len(X) * 0.7) -# 训练 model = LightGBMModel(params={"n_estimators": 200, "learning_rate": 0.03}) model.fit(X.iloc[:split], y.iloc[:split]) -# 回测 strategy = MLStrategy(model, fe) report = engine_bt.run(strategy, price, factor_df) -print(report.summary()) -print(model.get_feature_importance().head(5)) -``` - -### 示例 4:每日 Agent 运行 - -```python -from agents.orchestrator import AgentOrchestrator - -orch = AgentOrchestrator( - dm=dm, fe=engine_fe, bt=engine_bt, opt=opt_engine, sent=sent, -) -orch.setup() -results = orch.run_daily() - -# 获取结果 -sel = results["selection"] -risk = results["risk"] -report_path = results["report"]["report_path"] -print(f"日报: {report_path}") ``` --- @@ -893,268 +532,55 @@ print(f"日报: {report_path}") ### Q: SSH 隧道连接失败? ```bash -# 检查端口 lsof -i :13306 | grep LISTEN - -# 重新建立 bash shared/script/autossh.sh ``` ### Q: AkShare 返回 RemoteDisconnected? -这是 AkShare 的 curl_cffi 在连续请求时偶发的连接问题。系统已内置 3 次递增间隔重试 + fallback 机制,通常第 2-3 次重试会成功。如果持续失败: - -- 等待 30 秒后重试 -- 减少并发请求频率 -- 检查网络是否能访问 eastmoney.com +系统已内置 3 次递增间隔重试 + fallback 机制。如果持续失败:等待 30 秒后重试,或检查网络。 ### Q: 因子计算结果全是 NaN? - 技术因子:前 N 个周期内 NaN 是正常的(如 20 日动量前 19 天为 NaN) -- 基本面因子:检查财务数据是否已同步(`dm.get_financial(ts_code)`) +- 基本面因子:检查财务数据是否已同步 - 情绪因子:检查是否配置了 `QWEN_API_KEY` ### Q: 回测结果为 0 笔交易? -- 检查策略参数是否过于严格(如 RSI oversold=10 过少触发) +- 检查策略参数是否过于严格 - 使用 `OptunaEngine.optimize()` 寻找更优参数 -- 检查因子值是否合理(`factor_df.describe()`) ### Q: 模型训练只有 2 棵树? -当验证集损失不下降时,早停会在很少的迭代后触发。这是单股票预测的正常现象(信号噪声比低)。建议: - +单股票预测噪声比低,建议: - 设置 `eval_ratio=0.0` 禁用早停 - 降低 `learning_rate` 到 0.01 - 增加 `min_data_in_leaf` 防止过拟合 -### Q: 情绪因子返回空? - -- 确认 `.env` 中 `QWEN_API_KEY` 已配置 -- 检查网络是否能访问 `dashscope.aliyuncs.com` -- 如果使用本地 Ollama,确认服务运行中:`curl http://localhost:11434/api/tags` - ### Q: 日报中选股为空? -日报只对 DB 中有日线缓存的股票打分。需要先同步目标股票池的数据: - -```python -# 同步单只 -dm.sync_daily("000001.SZ") - -# 按范围批量同步(需先配置 SENTIMENT_SCOPE) -codes = sent.get_scope_stocks() -for code in codes[:10]: - dm.sync_daily(code) -``` - ---- - -## 13. CLI 脚本参考 - -所有脚本位于 `finance/cli/`,需在项目根目录或 `finance/` 下运行。 - ---- - -### 13.1 Agent 系统入口 — `agent_cli.py` +需要先同步目标股票池的数据: ```bash -cd finance && python cli/agent_cli.py <命令> [参数] +python finance/cli/agent_cli.py warmup 50 ``` -| 命令 | 说明 | 示例 | -|------|------|------| -| `daily [DATE]` | 完整每日流程(增量同步已缓存→评估风险→选股→日报) | `agent_cli.py daily` | -| `picks [N] [DATE]` | 多因子选股 Top N(需已缓存) | `agent_cli.py picks 15` | -| `risk` | 市场风险评估(等级、仓位、止损) | `agent_cli.py risk` | -| `research` | 因子发现:遍历因子计算 IC/IC_IR 排名 | `agent_cli.py research` | -| `report [DATE]` | 生成日报(含三指数行情+选股+情绪+风险评估) | `agent_cli.py report` | -| `warmup [N]` | 首次批量预热范围股票到 DB 缓存(每批 N 只,默认 50) | `agent_cli.py warmup 50` | - -**`daily` 流程**: - -``` -[Step 1/4] 增量同步 → 只更新已缓存股票(最新则 0.04s 跳过) - → 未缓存提示:运行 'agent_cli.py warmup' 首次预热 -[Step 2/4] 风险评估 → high/medium/low + 仓位建议 + 预警 -[Step 3/4] 股票打分 → DB 缓存命中率 + 多因子等权打分 → Top 15 -[Step 4/4] 日报生成 → 三指数行情 (Tushare) + 情绪摘要 + 风险预警 - → reports/daily_YYYYMMDD.md -``` - -**数据源优先级**:Tushare → AkShare(`.env` 配置 `TUSHARE_TOKEN`) - --- -### 13.2 数据层验证 — `demo_data_manager.py` +## 13. 文档索引 -```bash -python cli/demo_data_manager.py [--ts_code CODE] [--start YYYYMMDD] -``` - -| 参数 | 默认值 | 说明 | -|------|--------|------| -| `--ts_code` | `000001.SZ` | 测试股票代码 | -| `--start` | `20250101` | 起始日期 YYYYMMDD | - -5 步验证:数据库连接 → 建表 → 股票列表 → 日线获取(双源fallback) → 增量同步。 - ---- - -### 13.3 因子引擎验证 — `demo_factor_engine.py` - -```bash -python cli/demo_factor_engine.py [--ts_code CODE] [--ts_code2 CODE] -``` - -| 参数 | 默认值 | 说明 | -|------|--------|------| -| `--ts_code` | `000001.SZ` | 测试股票代码 | -| `--ts_code2` | `600519.SH` | 截面测试第二只股票 | - -验证:因子注册表(12分类/34因子)→ 技术因子计算(describe统计) → 基本面因子(ROE/PE/PB/EP) → NaN 覆盖率检查 → 双股票截面因子。 - ---- - -### 13.4 回测引擎验证 — `demo_backtest.py` - -```bash -python cli/demo_backtest.py [--ts_code CODE] -``` - -| 参数 | 默认值 | 说明 | -|------|--------|------| -| `--ts_code` | `000001.SZ` | 回测股票代码 | - -测试 5 个内置策略: - -| 策略 | 参数 | +| 文档 | 内容 | |------|------| -| SMACrossStrategy | (5,20) / (10,60) | -| RSIMeanRevertStrategy | (30,70) / (20,80) | -| MomentumBreakoutStrategy | lookback=20 | -| FactorCrossStrategy | momentum_20 > 0 | -| FactorRotationStrategy | momentum top 20% | - ---- - -### 13.5 参数优化验证 — `demo_optimizer.py` - -```bash -python cli/demo_optimizer.py [--ts_code CODE] [--trials N] -``` - -| 参数 | 默认值 | 说明 | -|------|--------|------| -| `--ts_code` | `000001.SZ` | 回测股票代码 | -| `--trials` | `200` | Optuna 试验次数 | - -对 RSI 反转策略执行参数寻优 + Walk-Forward 验证。输出最优 vs 默认对比表 + 参数重要性排序。 - ---- - -### 13.6 ML 模型验证 — `demo_ml.py` - -```bash -python cli/demo_ml.py [--ts_code CODE] [--lookahead N] -``` - -| 参数 | 默认值 | 说明 | -|------|--------|------| -| `--ts_code` | `000001.SZ` | 训练股票代码 | -| `--lookahead` | `5` | 预测未来 N 日收益 | - -完整 ML pipeline:特征工程(25因子→Winsorize→RobustScaler) → LightGBM训练(IC/CV) → CatBoost训练 → MLBenchmark对比(IC/收益/夏普/胜率)。 - ---- - -### 13.7 情绪因子快速验证 — `demo_sentiment.py` - -```bash -python cli/demo_sentiment.py [--ts_code CODE] [--no-qwen] -``` - -| 参数 | 默认值 | 说明 | -|------|--------|------| -| `--ts_code` | `000001.SZ` | 测试股票代码 | -| `--no-qwen` | flag | 跳过 Qwen API 调用 | - -6 步验证:新闻数据源(三源聚合) → 日期对齐 → Qwen 客户端状态 → SentimentEngine全链路 → 分析范围解析。 - -适合快速检查情绪因子系统是否就绪。 - ---- - -### 13.8 情绪因子详细演示 — `demo_sentiment_detail.py` - -```bash -python cli/demo_sentiment_detail.py [选项] -``` - -最详细的情绪因子脚本,支持完整命令行参数和逐步输出。 - -| 参数 | 类型 | 默认值 | 说明 | -|------|------|--------|------| -| `--ts_code` | str | `000001.SZ` | 股票代码,多个用逗号分隔 | -| `--date` | str | 今天 | 目标日期 YYYYMMDD | -| `--start` | str | date-30天 | 起始日期 YYYYMMDD | -| `--end` | str | date | 结束日期 YYYYMMDD | -| `--scope-type` | str | - | 分析范围:`index`/`sector`/`custom`/`all` | -| `--scope-indexes` | str | `000300` | 指数代码(逗号分隔) | -| `--scope-sectors` | str | - | 板块名称(逗号分隔) | -| `--max-news` | int | .env 配置 | 最大新闻条数 | -| `--max-analyze` | int | `50` | Qwen API 分析最大条数(控制成本) | -| `--no-xwlb` | flag | - | 禁用新闻联播数据源 | -| `--no-akshare` | flag | - | 禁用东方财富数据源 | -| `--no-mcp` | flag | - | 禁用 MCP 数据源 | -| `--source` | str | - | 仅用指定数据源:`xwlb`/`akshare`/`mcp` | -| `--no-qwen` | flag | - | 跳过 Qwen API 调用(仅演示数据流) | - -使用示例: - -```bash -# 默认演示(000001.SZ,最近30天,全数据源) -python cli/demo_sentiment_detail.py - -# 指定股票和日期 -python cli/demo_sentiment_detail.py --ts_code 600519.SH --date 20260603 - -# 多股票 + 日期范围 -python cli/demo_sentiment_detail.py --ts_code 000001.SZ,300316.SZ --start 20260501 --end 20260603 - -# 按指数成分股分析 -python cli/demo_sentiment_detail.py --scope-type index --scope-indexes 000300 - -# 按板块分析 -python cli/demo_sentiment_detail.py --scope-type sector --scope-sectors 银行,电力设备 - -# 只看东方财富新闻,不调用 Qwen -python cli/demo_sentiment_detail.py --source akshare --no-qwen --max-news 20 -``` - -输出 6 步详情: - -``` -Step 0: 初始化引擎(显示数据源、API状态、范围、Tushare可用性) -Step 1: 按数据源分别拉取新闻(xwlb/AkShare/MCP 各自数量 + 双源fallback) -Step 2: 新闻详情(按来源分开展示标题/内容/链接) -Step 3: 日期对齐(xwlb +1day偏移 + 非交易日对齐 + DB缓存检查→sync补齐) -Step 4: Qwen 情绪分析(每条新闻的分数/置信度/主题/来源标签) -Step 5: 因子计算(weighted sent / confidence-weighted / momentum + 公式说明) -Step 6: 结果输出(因子值表 + 历史统计 + 每新闻情绪贡献明细) -``` - ---- - -### 13.9 脚本一览 - -| 脚本 | 参数 | 用途 | 数据源 fallback | -|------|------|------|:---:| -| `agent_cli.py` | 子命令 + 参数 | 日常操作入口 | ✅ | -| `demo_data_manager.py` | `--ts_code` `--start` | Sprint 0 验证 | ✅ | -| `demo_factor_engine.py` | `--ts_code` `--ts_code2` | Sprint 1 验证 | ✅ | -| `demo_backtest.py` | `--ts_code` | Sprint 2 验证 | ✅ | -| `demo_optimizer.py` | `--ts_code` `--trials` | Sprint 3 验证 | ✅ | -| `demo_ml.py` | `--ts_code` `--lookahead` | Sprint 4 验证 | ✅ | -| `demo_sentiment.py` | `--ts_code` `--no-qwen` | Sprint 5 快速验证 | ✅ | -| `demo_sentiment_detail.py` | 14 个 argparse 参数 | Sprint 5 详细演示 | ✅ | +| [架构说明](architecture.md) | 项目架构、数据流、设计原则 | +| [开发指南](development.md) | 环境搭建、开发约定、模块说明 | +| [部署说明](deployment.md) | 本地环境、服务器、uWSGI、rsync 部署 | +| [因子与表结构速查](reference.md) | 34 因子注册表、DB 表结构、数据源接口 | +| [数据层详解](data-layer.md) | DataManager、数据库、缓存策略、已知 Bug | +| [因子引擎详解](factors.md) | 因子计算、情绪引擎、新闻源 | +| [回测引擎详解](backtest.md) | VectorBT、策略、信号工具、Optuna | +| [ML 模型详解](ml-models.md) | 特征工程、LightGBM/CatBoost、ML 策略 | +| [Agent 系统详解](agents.md) | Agent 架构、CLI、日报 | +| [DJAPI 接口](api.md) | Django API 端点参考 | +| [日报查询 API](news_report_api.md) | news/reports + news/events 接口 | +| [日报数据库](db_schema_v1.1.md) | news_report / news_event 表结构 | \ No newline at end of file diff --git a/init_plan.md b/init_plan.md deleted file mode 100644 index 029d4a4..0000000 --- a/init_plan.md +++ /dev/null @@ -1,675 +0,0 @@ -我认真看了你的目标和现有环境,我认为有一个关键点需要调整: - -**不要把 Claude Code Plugin 当成系统主体。** - -对于你的项目: - -```text -Claude Code -MCP-Hub -Django API -AkShare -MariaDB -VectorBT -Optuna -LightGBM -Qwen -``` - -Claude Code 应该只是: - -```text -AI开发助手 -AI研究助手 -``` - -而不是: - -```text -系统运行时核心 -``` - -真正的核心应该是: - -```text -finance/ -``` -finance 目录下python虚拟环境位于 finance/.venv/,在项目根目录下 可用 source finance/.venv/bin/activate激活 -这个目录未来即使你不用 Claude、换成 Cursor、Codex、OpenHands、Aider,都应该能独立运行。 - ---- - -# 推荐总体架构 - -未来你的根目录: - -```text -/Users/summer/Downloads/cc-cursor - -├── djapi/ -│ -├── finance/ -│ -├── mcp-servers/ -│ -├── shared/ -│ -├── docs/ -│ -└── .claude/ -``` - ---- - -# 各目录职责 - -## djapi - -仅负责: - -```text -数据库 -用户管理 -任务管理 -API接口 -报告管理 -``` - -类似: - -```text -Quant Platform Backend -``` - ---- - -## finance - -核心量化引擎 - -未来90%的代码都在这里。 - ---- - -结构: - -```text -finance/ - -├── config/ -│ -├── data/ -│ -├── factors/ -│ -├── models/ -│ -├── strategy/ -│ -├── optimizer/ -│ -├── backtest/ -│ -├── portfolio/ -│ -├── execution/ -│ -├── reports/ -│ -├── scheduler/ -│ -└── cli/ -``` - ---- - -# 第一阶段(V1) - -目标: - -```text -数据获取 -因子计算 -回测 -参数优化 -``` - -技术栈: - -```text -AkShare -MariaDB -VectorBT -Optuna -``` - ---- - -## 实施周期 - -### Week 1 - -基础设施 - ---- - -目录: - -```text -finance/ - -config/ -data/ -database/ -``` - ---- - -完成: - -### DataManager - -```python -class DataManager: -``` - -统一管理: - -```text -股票列表 -日线 -分钟线 -财务数据 -``` - ---- - -不要让策略直接调用: - -```python -ak.stock_zh_a_hist() -``` - -而是: - -```python -data_manager.get_daily() -``` - -这样以后换 TuShare 不改策略。 - ---- - -### Week 2 - -因子引擎 - -建立: - -```text -factors/ - -technical/ -fundamental/ -``` - ---- - -例如: - -```text -MomentumFactor -RSIFactor -ROEFactor -PEFactor -``` - -统一接口: - -```python -factor.calculate(df) -``` - ---- - -### Week 3 - -VectorBT回测层 - -建立: - -```text -backtest/vectorbt/ -``` - ---- - -统一接口: - -```python -engine.run(strategy) -``` - -以后: - -```text -VectorBT -Backtrader -Zipline -``` - -都能替换。 - ---- - -### Week 4 - -Optuna优化层 - -建立: - -```text -optimizer/ -``` - ---- - -统一接口: - -```python -optimizer.optimize(strategy) -``` - -实现: - -```text -MA -RSI -MACD -``` - -自动寻优。 - ---- - -# 第二阶段(V2) - -目标: - -```text -机器学习选股 -``` - -技术: - -```text -LightGBM -CatBoost -``` - ---- - -## 新增目录 - -```text -models/ - -├── lightgbm/ -└── catboost/ -``` - ---- - -建立统一模型接口: - -```python -class BaseModel: -``` - ---- - -例如: - -```python -fit() - -predict() - -save() - -load() -``` - ---- - -所有模型遵循: - -```python -BaseModel -``` - -接口。 - ---- - -# 第三阶段(V3) - -目标: - -```text -新闻因子 -公告因子 -研报因子 -``` - -技术: - -```text -Qwen -``` - ---- - -新增: - -```text -factors/sentiment/ - -models/qwen/ -``` - ---- - -Qwen职责: - -```text -文本转因子 -``` - -例如: - -```text -公告 -↓ -Qwen -↓ -sentiment_score -↓ -feature_101 -``` - ---- - -注意: - -Qwen不是策略。 - -Qwen是: - -```text -因子生产工具 -``` - ---- - -# 第四阶段(V4) - -Agent化 - -新增: - -```text -finance/agents/ -``` - ---- - -结构: - -```text -research_agent -selection_agent -risk_agent -report_agent -``` - ---- - -例如: - -### ResearchAgent - -负责: - -```text -发现新因子 -``` - ---- - -### SelectionAgent - -负责: - -```text -股票打分 -``` - ---- - -### RiskAgent - -负责: - -```text -仓位控制 -``` - ---- - -# Claude Code 接入时间点 - -很多人一开始就写 Skill。 - -我建议: - -### 不要现在写大量 Skill - -第一阶段只保留: - -```text -.claude/ - -skills/ - -factor-research -backtest -stock-selection -``` - -三个就够。 - ---- - -等 V2 完成以后再扩展。 - ---- - -# 数据流设计(必须遵守) - -未来所有代码都遵守: - -```text -Data - ↓ - -Factor - ↓ - -Model - ↓ - -Strategy - ↓ - -Backtest - ↓ - -Report -``` - ---- - -禁止: - -```text -Strategy - ↓ -直接访问AkShare -``` - ---- - -禁止: - -```text -Model - ↓ -直接访问数据库 -``` - ---- - -全部通过 Service 层。 - ---- - -# 推荐实施顺序 - -## Sprint 1(1~2周) - -完成: - -```text -finance/ -DataManager -MariaDB -AkShare -``` - -目标: - -```text -获取任意股票历史数据 -``` - ---- - -## Sprint 2(1周) - -完成: - -```text -Factor Engine -``` - -目标: - -```text -计算10个基础因子 -``` - ---- - -## Sprint 3(1周) - -完成: - -```text -VectorBT -``` - -目标: - -```text -回测单因子策略 -``` - ---- - -## Sprint 4(1周) - -完成: - -```text -Optuna -``` - -目标: - -```text -自动参数优化 -``` - ---- - -## Sprint 5(2周) - -完成: - -```text -LightGBM -CatBoost -``` - -目标: - -```text -训练未来5日收益预测模型 -``` - ---- - -## Sprint 6(2周) - -完成: - -```text -Qwen情绪因子 -``` - -目标: - -```text -新闻情绪打分 -``` - ---- - -## Sprint 7 - -完成: - -```text -Agent -Dashboard -自动日报 -``` - ---- - -按照这个路线,你的代码会从: - -```text -Mac Mini 单机量化研究平台 -``` - -平滑演进到: - -```text -多因子 + ML + LLM + Agent -量化研究平台 -``` - -中间不会出现“推倒重写”的情况。最关键的是先把 **DataManager → Factor Engine → Backtest Engine → Model Engine** 四个基础引擎设计好,后面的 LightGBM、Qwen、Agent 都只是插件式增加能力。 - diff --git a/reasonix.toml b/reasonix.toml deleted file mode 100644 index f10ff25..0000000 --- a/reasonix.toml +++ /dev/null @@ -1,5 +0,0 @@ -# 项目级 Reasonix 配置覆盖(仅此工作区生效) -# 本机未安装 bubblewrap,为执行运维命令关闭 bash 沙箱(全局配置仍为 enforce) -[sandbox] -bash = "off" -