Files
myquant/docs/ml-models.md
T
Simon 6acf938caf docs: 文档重构 — 清理 AI agent 残留,整合 docs/ 目录结构
- 删除 11 个残留文件: continuation.md, init_plan.md, reasonix.toml, djapi/continuation.md, djapi/.serena/, djapi/.claude/, djapi/.mcp.json, .claude/skills/, docs/usage.html, docs/db_schema.md, docs/report_db_design.md
- 7 个 CLAUDE-*.md 移入 docs/ 并重命名去 CLAUDE- 前缀
- 新增 4 个文档: architecture.md, development.md, api.md, deployment.md
- 重写 usage.md, README.md
- 修复所有过时引用和交叉链接
2026-08-22 11:56:40 +08:00

1.7 KiB
Raw Blame History

ML 模型层

FeatureEngine (finance/models/features.py)

因子 → 特征矩阵 + 标签。防前视偏差。

from models.features import FeatureEngine
fe = FeatureEngine(lookahead=5, label_type="regression")
X, y = fe.build(factor_df, price_df, fit=True)
# fit=True: Winsorize(1%/99%) → ffill → median fill → RobustScaler.fit → 标签计算
# fit=False: 复用训练时的 scaler + 有效特征

LightGBM (finance/models/lightgbm/model.py)

from models.lightgbm.model import LightGBMModel
model = LightGBMModel(params={"n_estimators": 200, "learning_rate": 0.03}, eval_ratio=0.2)
model.fit(X_train, y_train)     # val>=50 行才启用早停
pred = model.predict(X_test)
imp  = model.get_feature_importance()  # → DataFrame
cv   = model.cv_evaluate(X, y, n_folds=5)  # TimeSeriesSplit

CatBoost (finance/models/catboost/model.py)

同接口。get_feature_importance() / cv_evaluate()。

ML 策略 (finance/models/backtest_integration.py)

from models.backtest_integration import MLStrategy, MLBenchmark
strategy = MLStrategy(model, fe, buy_quantile=0.7, sell_quantile=0.3, rebalance_freq=5)
# 预测值分位 → 动态阈值 → 交易信号
benchmark = MLBenchmark([lgb, cb], fe, price_df, factor_df)
result = benchmark.run()  # → DataFrame: model × (IC, return, sharpe, win_rate, trades)

重要约束

  • lookahead 固定,不输入模型(防目标泄露)
  • 单股票 IC≈0 是正常现象(噪声主导),多股票截面才是 ML 发挥价值的地方
  • 特征工程严禁使用未来数据(RobustScaler fit 在训练集,transform 在测试集)
  • 交叉验证用 TimeSeriesSplit(不 shuffle)