修复:量价单位 / 未来函数守卫 / 实时画像闸门;行情回补到 2005;手册补全流程
本轮会话的三项正确性改造(均为「不报错、只让结果静默错」的类型):
1) 修复 stock_daily 量价单位前后不一致
- 现象:2015-2019 存 Tushare 原始单位(手/千元),2020 起存(股/元),2019 同日混合;
而流动性阈值按「元」配置 → 早年门槛实际是「日均成交额 ≥ 200 亿元」,
把 2015-2019 的股票池整体清空(实测 2016/2017/2018 各选出 0 只)。
- 修复:写入端 sync/price.py 统一换算;读取端 units.normalize_ohlcv_units
按行判定并幂等换算(price_history / avg_amount 都走它);
审计新增 UNIT-OHLCV 防回归。
- 效果:2016/2017/2018 的股票池变为 7/11/13 只。
2) 未来函数守卫(单次回测)
- 股票池自带 asof:若晚于回测起点即**拒绝执行**(原先静默冻结套用),
与 walk-forward 已有的拒绝理由一致;确需复现加 --allow-lookahead-universe,
偏差写入 unimplemented_json。
3) 新增实时(PIT)个股画像闸门
- profile/pit.py:每个决策日按当时可见数据重算过去 5 年画像,
惰性(仅买入条件已触发的标的)、面板按 asof 缓存、
规则不含财务指标时不查财报表;被剔除时产出 REJECT + 逐规则留痕。
- 指标定义复用 ProfileBuilder._profile_one(与批量画像逐值等价的回归测试)。
- profile/coverage.py:窗口覆盖率(按交易日历的真实开市天数),
策略新增 entry.profile_gate.min_window_coverage(默认 0,不改变既有行为)。
- core/metrics.py:闸门可用指标的唯一定义(配置期即校验,避免写错指标名静默失效)。
4) 行情回补到 2005(使 5/8/10 年窗口真正完整)
- stock_daily / adjust_factor / daily_basic 补到 2005-01-04;
hd_suspend / hd_limit 补到 2010-01-04。
- 5 年窗口覆盖率:2018-05-18 由 67.0% → 99.1%,2016-12-30 由 39.8% → 99.0%;
残差经逐日与 hd_suspend 交叉核实为真实停牌(16/16 命中)。
- 审计 G2/G3 与断点续传原先用固定阈值(2000 / 1500 只),
会把 2005-2009 的正常数据误判为异常 —— 改为按「当年应有上市股票数」成比例判定。
- 节流修正:daily/adj_factor/daily_basic 限频 480 → 170(实测该 token 约 196/min 即被拒)。
5) 自我声明如实化
- 原先「约束未生效」由「过滤后集合为空」判定,会把「这批股票恰好没停牌」
误报成「hd_suspend 无数据」;改为按表级判定。
- 补齐此前静默的「配置承诺但未实现」项:suspended_rule/limit_up_down_rule 的 defer、
cash_mode=reinvest/reinvest_rule、handle_rights_issue、signal_to_execution、
max_volume_pct、liquidity_limit_pct_adv —— 全部写入 unimplemented_json。
6) 手册:新增 §0「全流程操作(选股 → 画像 → 回测)」置于最前
- 逐步说明「命令做了什么、数据从哪来、落了哪些库、有哪些坑」;
含实时画像闸门 9 问 9 答、未来函数守卫表、成交与成本口径、验证 SQL。
- 修正旧 §2.4 漏传 --universe-run(选了池子却没用于回测);
修正两处声称「停牌顺延」「分红再投资」已实现的相反表述。
测试:403 项全部通过(含新增 test_units.py、test_profile_pit.py、
未实现声明诚实性测试、行序无关性回归测试)。
注意:本提交中 docs/*、README.md、src/hdiv/web/service.py 除本轮修改外,
也含此前遗留的未提交改动(无法按文件切分)。
This commit is contained in:
+256
-11
@@ -31,11 +31,11 @@ from hdiv.core.config import (
|
||||
config_hash,
|
||||
load_config,
|
||||
)
|
||||
from hdiv.core.errors import DataGapError
|
||||
from hdiv.core.errors import DataGapError, HdivError
|
||||
from hdiv.data import db
|
||||
from hdiv.data.repo import Repo, data_version
|
||||
from hdiv.data.sync.base import stable_id
|
||||
from hdiv.factor.dividend_yield import build_dps_events, ttm_dps_series
|
||||
from hdiv.factor.dividend_yield import build_dps_events, ttm_dps_series, ttm_params
|
||||
from hdiv.strategy.registry import StrategyRegistry
|
||||
|
||||
|
||||
@@ -160,6 +160,7 @@ class BacktestEngine:
|
||||
backtest: BacktestConfig | None = None,
|
||||
frozen_reference: tuple[date, date] | None = None,
|
||||
universe_run_id: str | None = None,
|
||||
allow_lookahead_universe: bool = False,
|
||||
) -> None:
|
||||
db.load_dotenv_once()
|
||||
self.strategy = strategy
|
||||
@@ -173,7 +174,14 @@ class BacktestEngine:
|
||||
# 指定 universe_run_id 时,使用该次筛选的成员作为**冻结股票池**
|
||||
# (不再按周期重新筛选)。这同时建立了「股票池记录 ↔ 回测记录」的显式关联。
|
||||
self.universe_run_id = universe_run_id
|
||||
# 股票池自带 asof:若它晚于回测起点,就等于用未来信息选股。
|
||||
# 默认拒绝;只有显式放行才执行,且必须把「含未来信息」写进 run 记录。
|
||||
self.allow_lookahead_universe = allow_lookahead_universe
|
||||
self.lookahead_universe_note: str | None = None
|
||||
self.universe_cfg = self.registry.resolved_universe(strategy)
|
||||
# 实时画像闸门(PIT):关闭时整条链路不参与,回测行为与启用前一致
|
||||
self.gate_cfg = strategy.entry.profile_gate
|
||||
self.pit: Any = None
|
||||
|
||||
@classmethod
|
||||
def from_strategy(cls, path: str | Path, **kw: Any) -> BacktestEngine:
|
||||
@@ -262,7 +270,25 @@ class BacktestEngine:
|
||||
bt.capital.initial,
|
||||
),
|
||||
"unimplemented": state["unimplemented"],
|
||||
"profile_gate": self.pit.stats() if self.pit is not None else None,
|
||||
}
|
||||
if self.pit is not None:
|
||||
gate_stats = self.pit.stats()
|
||||
result["profile_gate_verdicts"] = {
|
||||
"reject": sum(
|
||||
1 for x in state["signals"] if x.kind == "REJECT"
|
||||
),
|
||||
}
|
||||
if verbose:
|
||||
print(
|
||||
f" 实时画像:计算 {gate_stats['snapshots_computed']} 次"
|
||||
f"(缓存命中 {gate_stats['snapshots_cached']})"
|
||||
f",涉及 {gate_stats['distinct_asof']} 个决策时点"
|
||||
f",财报面板载入 {gate_stats['financial_loads']} 次"
|
||||
f",流动性查询 {gate_stats['liquidity_loads']} 次"
|
||||
f" | 画像剔除 {result['profile_gate_verdicts']['reject']} 次",
|
||||
flush=True,
|
||||
)
|
||||
if verbose:
|
||||
rc = result["reconciliation"]
|
||||
print(
|
||||
@@ -286,6 +312,55 @@ class BacktestEngine:
|
||||
)
|
||||
return result
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 股票池未来函数守卫
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
def _universe_asof(self) -> date | None:
|
||||
"""读股票池记录的 asof 日期(同时校验 run_id 是否存在)。"""
|
||||
df = db.read_sql(
|
||||
"SELECT asof_date FROM hd_universe_run WHERE run_id = :r",
|
||||
{"r": self.universe_run_id}, cfg=load_config("datasource"),
|
||||
)
|
||||
if df.empty:
|
||||
raise HdivError(
|
||||
f"股票池 {self.universe_run_id} 不存在(hd_universe_run 无此 run_id)。\n"
|
||||
f" 可执行 `python -m hdiv universe` 重新筛选,或在 Web 前端「股票池」页复制正确的 run_id。"
|
||||
)
|
||||
return pd.to_datetime(df["asof_date"].iloc[0]).date()
|
||||
|
||||
def _check_universe_asof(self, backtest_start: date) -> None:
|
||||
"""拒绝「用未来时点选出的股票池去跑更早的区间」。
|
||||
|
||||
这是本项目自己已经在 walk-forward 上认定的纪律(见
|
||||
``docs/implementation-status.md`` §7.4):股票池带 asof,把它套到更早的
|
||||
区间就是用未来信息选股。单次回测没有理由例外。
|
||||
"""
|
||||
uasof = self._universe_asof()
|
||||
if uasof is None or uasof <= backtest_start:
|
||||
return
|
||||
note = (
|
||||
f"股票池含未来信息:{self.universe_run_id} 的 asof={uasof} "
|
||||
f"晚于回测起点 {backtest_start},各调仓日复用了同一份事后名单"
|
||||
)
|
||||
if not self.allow_lookahead_universe:
|
||||
raise HdivError(
|
||||
f"拒绝执行:股票池的 asof({uasof})晚于回测起点({backtest_start})。\n"
|
||||
f" 股票池 {self.universe_run_id} 是用 {uasof} 当天可见的数据选出来的,\n"
|
||||
f" 名单里含有回测起点时不可能知道的信息(哪些公司此后仍满足连续分红、\n"
|
||||
f" 5 年 ROE、自由现金流覆盖等条件)。把它套到更早的年份即未来函数。\n"
|
||||
f" 正确做法(推荐第 1 种):\n"
|
||||
f" 1) 去掉 --universe-run,让引擎在每个调仓日按当时可见数据重新筛选:\n"
|
||||
f" python -m hdiv backtest --start {backtest_start} --end <end>\n"
|
||||
f" 2) 若只想检验某个固定股票池,把回测起点改到该股票池 asof 之后:\n"
|
||||
f" python -m hdiv backtest --universe-run {self.universe_run_id} "
|
||||
f"--start {uasof}\n"
|
||||
f" 确实需要复现「带未来信息」的历史结果(例如与修复前的记录对比)时,\n"
|
||||
f" 显式放行:--allow-lookahead-universe\n"
|
||||
f" 届时本次运行会在 hd_backtest_run.unimplemented_json 中如实声明该偏差。"
|
||||
)
|
||||
self.lookahead_universe_note = note
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 数据准备
|
||||
# ------------------------------------------------------------------
|
||||
@@ -310,6 +385,11 @@ class BacktestEngine:
|
||||
del cur
|
||||
universe_by_refresh: dict[date, set[str]] = {}
|
||||
if self.universe_run_id:
|
||||
# 未来函数守卫:股票池自带 asof。若它晚于回测起点,名单里就含有
|
||||
# 「当时不可能知道」的信息(哪些公司此后仍满足分红/质量条件),
|
||||
# 把它套到更早的年份上就是用未来信息选股 —— 与 walk-forward
|
||||
# 拒绝 --universe-run 是同一条理由,这里必须同样拒绝。
|
||||
self._check_universe_asof(days[0])
|
||||
# 冻结股票池:直接取该次筛选的入选成员,所有调仓日复用同一份清单。
|
||||
# 好处是可复现(同一 run_id 永远对应同一股票池),并建立双向关联。
|
||||
dfu = db.read_sql(
|
||||
@@ -366,6 +446,24 @@ class BacktestEngine:
|
||||
suspend = self._load_suspend(all_syms, days[0], days[-1])
|
||||
limits = self._load_limits(all_syms, days[0], days[-1])
|
||||
|
||||
# --- 实时画像闸门:预载跨决策日共享的面板(仅在启用时)---
|
||||
if self.gate_cfg.enabled:
|
||||
from hdiv.profile.pit import PitProfileService
|
||||
|
||||
self.pit = PitProfileService(window_years=self.gate_cfg.window_years)
|
||||
# 画像取数起点必须覆盖最长窗口(profile.yml 的 windows_years),
|
||||
# 与分位参照窗口(backtest.yml 的 lookback_years)是两个独立的量。
|
||||
pit_start = date(max(days[0].year - self.pit.max_years - 1, 2000), 1, 1)
|
||||
self.pit.prepare(all_syms, pit_start, days[-1])
|
||||
self.pit.configure({r.metric for r in self.gate_cfg.rules})
|
||||
if verbose:
|
||||
print(
|
||||
f" 实时画像闸门已启用:窗口 {self.gate_cfg.window_years} 年,"
|
||||
f"{len(self.gate_cfg.rules)} 条规则,"
|
||||
f"面板自 {pit_start} 起载入({len(all_syms)} 只)",
|
||||
flush=True,
|
||||
)
|
||||
|
||||
return {
|
||||
"days": days,
|
||||
"refresh_dates": refresh_dates,
|
||||
@@ -421,6 +519,16 @@ class BacktestEngine:
|
||||
)
|
||||
return out
|
||||
|
||||
def _has_constraint_rows(self, table: str, start: date, end: date) -> bool:
|
||||
"""该约束表在回测区间内**是否有任何数据**(表级判定,与股票池无关)。"""
|
||||
if not db.table_exists(table, self.cfg_db()):
|
||||
return False
|
||||
df = db.read_sql(
|
||||
f"SELECT COUNT(*) AS n FROM {table} WHERE trade_date BETWEEN :s AND :e",
|
||||
{"s": start, "e": end}, cfg=self.cfg_db(),
|
||||
)
|
||||
return bool(df["n"].iloc[0])
|
||||
|
||||
def cfg_db(self) -> Any:
|
||||
return load_config("datasource")
|
||||
|
||||
@@ -527,13 +635,54 @@ class BacktestEngine:
|
||||
peak = equity["total_value"].cummax()
|
||||
equity["drawdown"] = equity["total_value"] / peak - 1.0
|
||||
|
||||
for name in ("涨跌停近似", "停牌顺延", "成交量占比约束"):
|
||||
if name == "涨跌停近似" and not ctx["limits"]:
|
||||
unimplemented.add("涨跌停约束未生效(hd_limit 无数据,成交按可达价格近似)")
|
||||
if name == "停牌顺延" and not ctx["suspend"]:
|
||||
unimplemented.add("停牌约束未生效(hd_suspend 无数据)")
|
||||
# 「约束没生效」只能由**表里有没有数据**判定,不能由「过滤后集合为空」判定 ——
|
||||
# 后者会把「这批股票这段时间恰好没停牌/没涨跌停」误报成「数据缺失」。
|
||||
# 实测:2026-08~09 的 hd_suspend 覆盖到 2026-09-30,却因该区间无停牌
|
||||
# 而被声明成「hd_suspend 无数据」,属于把自己的建模正常状态说成数据缺陷。
|
||||
if not self._has_constraint_rows("hd_limit", days[0], days[-1]):
|
||||
unimplemented.add("涨跌停约束未生效(hd_limit 在回测区间内无数据,成交按可达价格近似)")
|
||||
if not self._has_constraint_rows("hd_suspend", days[0], days[-1]):
|
||||
unimplemented.add("停牌约束未生效(hd_suspend 在回测区间内无数据)")
|
||||
# 以下三项**配置写了但引擎没实现**,必须如实声明 —— 否则 run 记录看起来
|
||||
# 「一切正常」,而使用者以为 backtest.yml 的 defer / reinvest 生效了。
|
||||
# (配置承诺与实际行为不一致,是本项目反复记录的一类缺陷。)
|
||||
# 注意字段归属:fill/dividend 在 backtest.yml;execution/risk 在策略 yml。
|
||||
if str(bt.fill.suspended_rule) != "skip" or str(bt.fill.limit_up_down_rule) != "skip" \
|
||||
or str(s.execution.suspended_rule) != "skip" \
|
||||
or str(s.execution.limit_up_down_rule) != "skip":
|
||||
unimplemented.add(
|
||||
"未实现停牌/涨跌停顺延(suspended_rule / limit_up_down_rule 的 "
|
||||
"defer 分支):未成交信号在当日被**丢弃**,不会顺延到下一个可成交日"
|
||||
)
|
||||
if str(bt.dividend.cash_mode) != "hold" or bt.dividend.reinvest_rule:
|
||||
unimplemented.add(
|
||||
"未实现分红再投资规则(cash_mode=reinvest / reinvest_rule):"
|
||||
"现金分红按除权日入账后**留存为现金**,在下次调仓时按目标权重重新配置"
|
||||
)
|
||||
if s.execution.signal_to_execution != "next_open" or str(bt.fill.price) != "next_open":
|
||||
unimplemented.add(
|
||||
f"未实现 signal_to_execution/fill.price 的 "
|
||||
f"{s.execution.signal_to_execution}/{bt.fill.price} 分支:"
|
||||
f"成交固定按信号次日开盘价"
|
||||
)
|
||||
if bt.dividend.handle_rights_issue:
|
||||
unimplemented.add(
|
||||
"未实现配股处理(handle_rights_issue):配股缴款/股数变动不入账"
|
||||
)
|
||||
if not bt.fill.partial_fill:
|
||||
unimplemented.add("未启用部分成交(按信号全额成交,但受资金与权重上限约束)")
|
||||
if bt.fill.max_volume_pct is not None:
|
||||
unimplemented.add(
|
||||
"未实现成交量占比约束(fill.max_volume_pct 未被使用)"
|
||||
)
|
||||
if s.risk.liquidity_limit_pct_adv is not None:
|
||||
unimplemented.add(
|
||||
"未实现 risk.liquidity_limit_pct_adv(单笔成交不超过当日成交额的比例)"
|
||||
)
|
||||
# 显式放行的未来函数必须留在 run 记录里 —— 否则事后无法分辨
|
||||
# 「这条收益曲线是干净的」还是「这条用了事后名单」。
|
||||
if self.lookahead_universe_note:
|
||||
unimplemented.add(self.lookahead_universe_note)
|
||||
|
||||
div_df = pd.DataFrame(dividend_ledger)
|
||||
return {
|
||||
@@ -574,9 +723,12 @@ class BacktestEngine:
|
||||
if close_hist.empty:
|
||||
continue
|
||||
idx = pd.DatetimeIndex(close_hist.index)
|
||||
# 参数从因子层的统一来源取,不再硬编码 ——
|
||||
# 否则改了 profile.yml 的回测也不会变(曾如此)。
|
||||
_w, _g, _sm = ttm_params()
|
||||
dps = ttm_dps_series(
|
||||
idx, ctx["events"].get(sym, pd.DataFrame()),
|
||||
ttm_days=365, grace_days=45,
|
||||
ttm_days=_w, grace_days=_g, smooth_spikes=_sm,
|
||||
)
|
||||
with np.errstate(divide="ignore", invalid="ignore"):
|
||||
y = np.where(close_hist.to_numpy(dtype="float64") > 0,
|
||||
@@ -589,6 +741,14 @@ class BacktestEngine:
|
||||
ref_ser = ser.loc[pd.Timestamp(ref[0]): pd.Timestamp(ref[1])] if ref else ser
|
||||
if ref_ser.empty:
|
||||
ref_ser = ser
|
||||
# 样本不足则不作判断 —— 保持现有仓位,既不买也不卖。
|
||||
#
|
||||
# 分位 = 「≤当前值的观测占比」。窗口只有 1 个观测且恰好等于当前值时
|
||||
# 占比 100%,会击穿任何买入阈值。这是统计假象而非「股息率处于高位」:
|
||||
# 实测 2015-01-06(行情数据首日)窗口 2010-2015 只有 1 个观测,
|
||||
# 8 只股票因此被「100% 分位」买入。
|
||||
if ref_ser.size < self.bt_cfg.percentile_reference.min_observations:
|
||||
continue
|
||||
pct = float((ref_ser <= current).sum() / ref_ser.size * 100.0)
|
||||
|
||||
held = sym in positions
|
||||
@@ -601,6 +761,8 @@ class BacktestEngine:
|
||||
"reference_window": [str(ref[0]), str(ref[1])] if ref else None,
|
||||
"reference_mode": self._reference_mode(),
|
||||
"observation_count": int(ref_ser.size),
|
||||
"min_observations": int(
|
||||
self.bt_cfg.percentile_reference.min_observations),
|
||||
"close": price,
|
||||
}
|
||||
|
||||
@@ -611,12 +773,19 @@ class BacktestEngine:
|
||||
|
||||
if not held:
|
||||
if target is not None and target > 0 and pct >= s.entry.yield_percentile:
|
||||
gate = self._gate(sym, day)
|
||||
if gate is not None and gate["verdict"] != "PASS":
|
||||
out.append(self._reject_signal(
|
||||
sym, day, current, pct, price, common, gate,
|
||||
))
|
||||
continue
|
||||
out.append(Signal(
|
||||
sym, day, "BUY", target, current, pct, price,
|
||||
{**common,
|
||||
"rule": f"股息率历史分位 {pct:.1f}% >= P{s.entry.yield_percentile:g},"
|
||||
f"目标仓位 {target:.0%}",
|
||||
"reason_cn": "股息率进入历史高位区间,达到买入阈值"},
|
||||
"reason_cn": "股息率进入历史高位区间,达到买入阈值",
|
||||
**({"profile_gate": gate} if gate else {})},
|
||||
))
|
||||
else:
|
||||
if target is None:
|
||||
@@ -628,15 +797,91 @@ class BacktestEngine:
|
||||
"rule": f"股息率历史分位 {pct:.1f}% <= P{s.exit.yield_percentile:g}",
|
||||
"reason_cn": "股息率回落至历史低位区间,达到卖出阈值,清仓"},
|
||||
))
|
||||
else:
|
||||
elif target < 1.0:
|
||||
out.append(Signal(
|
||||
sym, day, "TRIM" if target < 1.0 else "ADD", target, current, pct, price,
|
||||
sym, day, "TRIM", target, current, pct, price,
|
||||
{**common,
|
||||
"rule": f"分位 {pct:.1f}% 对应目标仓位 {target:.0%}",
|
||||
"reason_cn": "股息率分位变动,按阶梯规则调整仓位"},
|
||||
))
|
||||
else:
|
||||
# ADD 也是买入 —— 同样要过实时画像闸门。
|
||||
# 被拒时**不动已有仓位**(REJECT 不进入待成交队列),
|
||||
# 因为闸门的语义是「不值得买」,不是「该卖」。
|
||||
gate = self._gate(sym, day)
|
||||
if gate is not None and gate["verdict"] != "PASS":
|
||||
out.append(self._reject_signal(
|
||||
sym, day, current, pct, price, common, gate,
|
||||
))
|
||||
continue
|
||||
out.append(Signal(
|
||||
sym, day, "ADD", target, current, pct, price,
|
||||
{**common,
|
||||
"rule": f"分位 {pct:.1f}% 对应目标仓位 {target:.0%}",
|
||||
"reason_cn": "股息率分位变动,按阶梯规则调整仓位",
|
||||
**({"profile_gate": gate} if gate else {})},
|
||||
))
|
||||
return out
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 实时画像闸门
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
def _gate(self, sym: str, day: date) -> dict[str, Any] | None:
|
||||
"""惰性计算该股在 ``day`` 的实时画像并判定闸门。
|
||||
|
||||
只在「买入条件已触发」时调用 —— 这是「在触发条件的时候计算」的落点。
|
||||
未启用闸门时返回 ``None``,调用方不产生任何额外行为。
|
||||
|
||||
**启用但未初始化必须报错,不得静默放行**:静默放行等于
|
||||
「配置了一个风险控制但它不生效」,属于最难发现的一类失效 ——
|
||||
回测照跑,拿到的却是没有闸门的版本。
|
||||
"""
|
||||
if not self.gate_cfg.enabled:
|
||||
return None
|
||||
if self.pit is None:
|
||||
raise HdivError(
|
||||
"实时画像闸门已启用,但画像服务未初始化。\n"
|
||||
" 正常路径由 BacktestEngine.run() → _prepare() 负责初始化;\n"
|
||||
" 若你直接调用 _evaluate/_simulate,请先调用 _prepare(days)。"
|
||||
)
|
||||
from hdiv.profile.pit import evaluate_gate
|
||||
|
||||
snap = self.pit.snapshot(sym, day)
|
||||
rules = [r.model_dump() for r in self.gate_cfg.rules]
|
||||
return evaluate_gate(
|
||||
rules, snap,
|
||||
on_unverifiable=self.gate_cfg.on_unverifiable,
|
||||
min_window_coverage=self.gate_cfg.min_window_coverage,
|
||||
)
|
||||
|
||||
def _reject_signal(
|
||||
self, sym: str, day: date, current: float, pct: float, price: float,
|
||||
common: dict[str, Any], gate: dict[str, Any],
|
||||
) -> Signal:
|
||||
"""把「为什么不买」写成可追溯的信号记录(plan.md §32 的同一条原则)。"""
|
||||
failed = gate.get("failed") or []
|
||||
unver = gate.get("unverifiable") or []
|
||||
if failed:
|
||||
why = "实时画像未通过:" + "、".join(failed)
|
||||
cn = "按当日可见数据重算画像后判定不值得买,剔除"
|
||||
else:
|
||||
why = "实时画像无法验证:" + "、".join(unver)
|
||||
cn = "按当日可见数据重算画像,样本不足/数据缺失,保守不买"
|
||||
checks = {
|
||||
c["metric"]: {"stat": c["stat"], "actual": c["actual"],
|
||||
"status": c["status"], "threshold": c["threshold"],
|
||||
"op": c["op"], "passed": c["passed"]}
|
||||
for c in gate.get("checks", [])
|
||||
}
|
||||
return Signal(
|
||||
sym, day, "REJECT", 0.0, current, pct, price,
|
||||
{**common, "rule": why, "reason_cn": cn,
|
||||
"profile_gate": gate, "profile_checks": checks,
|
||||
# executed=False + skip_reason 让它在「未成交信号」列表里可读
|
||||
"skip_reason": "PROFILE_GATE", "executed": False},
|
||||
)
|
||||
|
||||
def _reference_window(self, day: date) -> tuple[date, date] | None:
|
||||
if self.frozen_reference is not None:
|
||||
return self.frozen_reference
|
||||
|
||||
@@ -27,6 +27,8 @@ import numpy as np
|
||||
import pandas as pd
|
||||
|
||||
from hdiv.core.config import BacktestConfig, load_config
|
||||
|
||||
from hdiv.report.format import NumFmt
|
||||
from hdiv.core.errors import DataGapError, SchemaValidationError
|
||||
from hdiv.data import db
|
||||
from hdiv.data.repo import Repo, data_version
|
||||
@@ -222,6 +224,7 @@ class WalkForwardRunner:
|
||||
from hdiv.factor.dividend_yield import (
|
||||
build_dps_events,
|
||||
dividend_yield_series,
|
||||
ttm_params,
|
||||
)
|
||||
|
||||
sel = self.registry.selector(self.strategy)
|
||||
@@ -236,8 +239,10 @@ class WalkForwardRunner:
|
||||
for sym, g in px.groupby("symbol"):
|
||||
g = g.sort_values("trade_date").copy()
|
||||
g["trade_date"] = pd.to_datetime(g["trade_date"])
|
||||
_w, _g, _sm = ttm_params()
|
||||
ser = dividend_yield_series(
|
||||
g.set_index("trade_date")["close"], events.get(sym, pd.DataFrame())
|
||||
g.set_index("trade_date")["close"], events.get(sym, pd.DataFrame()),
|
||||
ttm_days=_w, grace_days=_g, smooth_spikes=_sm,
|
||||
)
|
||||
if ser.empty:
|
||||
continue
|
||||
@@ -315,7 +320,7 @@ class WalkForwardRunner:
|
||||
|
||||
def pct(k: str) -> str:
|
||||
v = s.get(k)
|
||||
return "—" if v is None else f"{v * 100:,.2f}%"
|
||||
return "—" if v is None else NumFmt.from_config().pct(v)
|
||||
|
||||
def num(k: str, d: int = 2) -> str:
|
||||
v = s.get(k)
|
||||
|
||||
+46
-4
@@ -104,10 +104,15 @@ def cmd_sync(args: argparse.Namespace) -> int:
|
||||
elif target == "backfill":
|
||||
from hdiv.data.sync import price
|
||||
|
||||
# basic_start 必须可传入:run_backfill 的默认值是 2015-01-01,
|
||||
# 若要回补 2015 年之前的 daily_basic,只给 --basic-end 是不够的 ——
|
||||
# 早期 CLI 没有这个参数,于是「回补 2010-2014」实际只补了行情,
|
||||
# daily_basic 仍停在 2015,画像里的 PE/PB 依旧拿不到早年数据。
|
||||
r = price.run_backfill(
|
||||
price_start=args.start,
|
||||
price_end=args.end,
|
||||
basic_end=args.basic_end,
|
||||
basic_start=args.basic_start or args.start,
|
||||
basic_end=args.basic_end or args.end,
|
||||
resume=not args.no_resume,
|
||||
)
|
||||
return 0 if r.get("monotonic_ok", True) else 1
|
||||
@@ -234,6 +239,13 @@ def cmd_profile(args: argparse.Namespace) -> int:
|
||||
pb = ProfileBuilder.from_config(args.config)
|
||||
result = pb.run(universe_run_id=args.universe_run, symbols=args.symbols, asof=args.asof)
|
||||
print(f"画像 {result['run_id']}:{result['symbol_count']} 只")
|
||||
# 覆盖率必须显眼:名义「5 年窗口」可能只有 3 年多数据(数据起点截断)
|
||||
for w in result.get("warnings", []):
|
||||
print(f" ⚠ {w}")
|
||||
worst = (result.get("window_coverage") or {}).get("worst")
|
||||
if worst:
|
||||
wy, cov = worst
|
||||
print(f" 窗口覆盖率:最差 {wy} 年窗口 {cov:.1%}(1.0 = 名义窗口被完整覆盖)")
|
||||
if args.html:
|
||||
from hdiv.report.build import build_profile_report
|
||||
|
||||
@@ -278,6 +290,23 @@ def cmd_backtest(args: argparse.Namespace) -> int:
|
||||
from hdiv.backtest.walk_forward import WalkForwardRunner
|
||||
|
||||
if args.mode == "walkforward":
|
||||
# 冻结股票池与 walk-forward 在时序上不兼容:
|
||||
# 股票池有其自身的 asof(如 2025-01-21),而 walk-forward 的窗口从
|
||||
# 2015 年就开始训练 —— 用未来时点选出的股票池去跑过去的窗口,
|
||||
# 就是典型的未来函数,恰好破坏 walk-forward 要守护的纪律。
|
||||
#
|
||||
# 早期实现没有这个参数,于是 --universe-run 被**静默忽略**,
|
||||
# 用户以为按自己的股票池跑了,实际跑的是逐窗口自筛选。
|
||||
if args.universe_run:
|
||||
raise HdivError(
|
||||
"walk-forward 模式不支持 --universe-run。\n"
|
||||
" 原因:股票池带有自己的时点(asof),把它套到更早的训练窗口上\n"
|
||||
" 等于用未来信息选股,会破坏 walk-forward 的无未来函数纪律。\n"
|
||||
" 正确做法:walk-forward 会在每个窗口内按各自时点重新筛选,\n"
|
||||
" 这正是它要检验的「策略能否在未知未来上复现」。\n"
|
||||
" 若确实想检验某个固定股票池,请用普通回测:\n"
|
||||
" python -m hdiv backtest --universe-run <run_id>"
|
||||
)
|
||||
wf = WalkForwardRunner.from_strategy(args.strategy)
|
||||
res = wf.run()
|
||||
print(f"Walk-forward {res['wf_id']}:{res['window_count']} 个窗口")
|
||||
@@ -289,7 +318,9 @@ def cmd_backtest(args: argparse.Namespace) -> int:
|
||||
|
||||
_reject_no_persist_with_html(args, "hdiv backtest")
|
||||
engine = BacktestEngine.from_strategy(
|
||||
args.strategy, universe_run_id=args.universe_run
|
||||
args.strategy,
|
||||
universe_run_id=args.universe_run,
|
||||
allow_lookahead_universe=args.allow_lookahead_universe,
|
||||
)
|
||||
res = engine.run(start=args.start, end=args.end, persist=not args.no_persist)
|
||||
if args.no_persist:
|
||||
@@ -366,7 +397,12 @@ def build_parser() -> argparse.ArgumentParser:
|
||||
s.add_argument("--which", choices=["daily", "adj_factor", "daily_basic"])
|
||||
s.add_argument("--start", default="2015-01-01")
|
||||
s.add_argument("--end", default="2018-12-31")
|
||||
s.add_argument("--basic-end", default="2019-12-31")
|
||||
s.add_argument(
|
||||
"--basic-start", default=None,
|
||||
help="daily_basic 回补起点(默认与 --start 相同)。"
|
||||
"回补 2015 年之前的数据时必须显式给出,否则 daily_basic 不会被回补",
|
||||
)
|
||||
s.add_argument("--basic-end", default="2019-12-31", help="daily_basic 回补终点")
|
||||
s.add_argument("--no-resume", action="store_true")
|
||||
s.add_argument("--no-weight", action="store_true")
|
||||
s.set_defaults(func=cmd_sync)
|
||||
@@ -477,7 +513,13 @@ def build_parser() -> argparse.ArgumentParser:
|
||||
b.add_argument("--end", default=None)
|
||||
b.add_argument(
|
||||
"--universe-run", default=None,
|
||||
help="使用指定股票池筛选记录的成员作为冻结股票池(并建立关联)",
|
||||
help="使用指定股票池筛选记录的成员作为冻结股票池(并建立关联);"
|
||||
"若该股票池的 asof 晚于回测起点则拒绝执行(未来函数)",
|
||||
)
|
||||
b.add_argument(
|
||||
"--allow-lookahead-universe", action="store_true",
|
||||
help="显式放行「股票池 asof 晚于回测起点」的组合(含未来信息,"
|
||||
"会如实写入 hd_backtest_run.unimplemented_json)",
|
||||
)
|
||||
b.add_argument("--no-persist", action="store_true")
|
||||
# HTML 报告已降级为「导出件」:默认不生成,需要时显式 --html。
|
||||
|
||||
@@ -96,11 +96,36 @@ class PathsConfig(StrictModel):
|
||||
log_dir: str = "logs"
|
||||
|
||||
|
||||
class SyncConfig(StrictModel):
|
||||
"""行情同步的「完整性」判定口径(决定断点续传会不会重拉整段历史)。
|
||||
|
||||
一个交易日被视为**已完整同步**,要求当日股票数 ≥
|
||||
``max(min_symbols_floor, min_symbols_ratio × 当年应有上市股票数)``。
|
||||
|
||||
**为什么不能只用一个绝对阈值**:A 股 2005 年只有约 1,350 只股票,
|
||||
2010 年约 1,700 只。若固定要求 1,500 只,2005-2009 的**每一个交易日**
|
||||
都会被判成「未完成」,于是断点续传完全失效 —— 回补中断一次就要从
|
||||
第一天重新拉,且每次重跑都会把整段早年历史再拉一遍。
|
||||
"""
|
||||
|
||||
min_symbols_floor: int = 200
|
||||
min_symbols_ratio: float = 0.6
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _check(self) -> SyncConfig:
|
||||
if self.min_symbols_floor < 1:
|
||||
raise SchemaValidationError("sync.min_symbols_floor 必须为正")
|
||||
if not 0.0 < self.min_symbols_ratio <= 1.0:
|
||||
raise SchemaValidationError("sync.min_symbols_ratio 必须落在 (0, 1]")
|
||||
return self
|
||||
|
||||
|
||||
class DataSourceConfig(StrictModel):
|
||||
version: int = 1
|
||||
database: DatabaseConfig
|
||||
tushare: TushareConfig = Field(default_factory=TushareConfig)
|
||||
paths: PathsConfig = Field(default_factory=PathsConfig)
|
||||
sync: SyncConfig = Field(default_factory=SyncConfig)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -292,6 +317,8 @@ class SufficiencyConfig(StrictModel):
|
||||
class TtmDividendConfig(StrictModel):
|
||||
window_days: int = 365
|
||||
grace_days: int = 45
|
||||
# 是否消除除权间隔不规整造成的毛刺(重叠虚高 / 断档虚低)
|
||||
smooth_spikes: bool = True
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _check(self) -> TtmDividendConfig:
|
||||
@@ -479,6 +506,7 @@ class ScheduleConfig(StrictModel):
|
||||
class PercentileReferenceConfig(StrictModel):
|
||||
mode: Literal["rolling", "frozen"] = "rolling"
|
||||
lookback_years: int = 5
|
||||
min_observations: int = 250
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _check(self) -> PercentileReferenceConfig:
|
||||
@@ -602,11 +630,84 @@ class ScaleStep(StrictModel):
|
||||
return self
|
||||
|
||||
|
||||
class ProfileGateRule(StrictModel):
|
||||
"""一条实时画像闸门规则。
|
||||
|
||||
语义:``<metric> 的 <stat> <op> <value>`` 必须成立,否则不买。
|
||||
指标名与分位可用性在**配置期**校验 —— 写错一个指标名若拖到运行时,
|
||||
只会得到「无法验证 → 保守不买」,表现为策略再也不交易,极难定位。
|
||||
"""
|
||||
|
||||
metric: str
|
||||
stat: Literal["current_value", "current_percentile"] = "current_value"
|
||||
op: Literal[">=", "<=", ">", "<"] = ">="
|
||||
value: float
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _check(self) -> ProfileGateRule:
|
||||
from hdiv.core.metrics import GATE_METRICS, PERCENTILE_METRICS
|
||||
|
||||
if self.metric not in GATE_METRICS:
|
||||
head = self.metric.split("_")[0]
|
||||
near = sorted(m for m in GATE_METRICS if head and head in m)
|
||||
raise SchemaValidationError(
|
||||
f"profile_gate 规则引用了未知指标 {self.metric!r}。"
|
||||
f"可选指标见 hdiv/core/metrics.py;相近的有 {near[:6]}"
|
||||
)
|
||||
if self.stat == "current_percentile" and self.metric not in PERCENTILE_METRICS:
|
||||
raise SchemaValidationError(
|
||||
f"{self.metric} 是标量指标,没有历史分位,不能用 "
|
||||
f"stat=current_percentile。有分位的指标:{sorted(PERCENTILE_METRICS)}"
|
||||
)
|
||||
return self
|
||||
|
||||
|
||||
class ProfileGateConfig(StrictModel):
|
||||
"""实时(PIT)个股画像闸门。
|
||||
|
||||
在每个决策日、**买入条件已经触发之后**,用「当时可见的数据」重算画像,
|
||||
不通过的票直接剔除。跨股票共享的面板按时点缓存,代价与「触发次数」成正比,
|
||||
而不是与「回测区间 × 股票数」成正比。
|
||||
"""
|
||||
|
||||
#: 是否启用。关闭时回测行为与启用前完全一致(可用于复现历史结果)
|
||||
enabled: bool = False
|
||||
#: 画像统计窗口(年)。0 = 全历史;其余必须是 config/profile.yml 的 windows_years 之一
|
||||
window_years: int = 5
|
||||
#: 数据缺失/样本不足(无法验证)时:reject = 保守不买,pass = 放行
|
||||
on_unverifiable: Literal["reject", "pass"] = "reject"
|
||||
#: 窗口**实际覆盖率**下限(1.0 = 名义 5 年就必须真有 5 年数据)。
|
||||
#: 0 = 不因覆盖率淘汰(默认,保持改造前行为)。
|
||||
#:
|
||||
#: 为什么需要它:`window_slice(asof, 5)` 只是「把已有数据切成最近 5 年」,
|
||||
#: 数据起点晚于窗口左端时窗口会被静默截短 —— 实测 2018-05-18 的「5 年」
|
||||
#: 窗口只有 3.4 年(817/1215 个交易日,67%),而画像仍报 OK。
|
||||
min_window_coverage: float = 0.0
|
||||
rules: list[ProfileGateRule] = Field(default_factory=list)
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _check(self) -> ProfileGateConfig:
|
||||
if self.window_years < 0:
|
||||
raise SchemaValidationError("profile_gate.window_years 不能为负")
|
||||
if not 0.0 <= self.min_window_coverage <= 1.0:
|
||||
raise SchemaValidationError(
|
||||
f"profile_gate.min_window_coverage 必须落在 [0, 1],"
|
||||
f"当前 {self.min_window_coverage}"
|
||||
)
|
||||
if self.enabled and not self.rules:
|
||||
raise SchemaValidationError(
|
||||
"profile_gate.enabled=true 但 rules 为空 —— 空闸门等于每次都要"
|
||||
"算一遍画像再无条件放行。请补齐规则,或把 enabled 设为 false。"
|
||||
)
|
||||
return self
|
||||
|
||||
|
||||
class EntryConfig(StrictModel):
|
||||
yield_percentile: float
|
||||
require_universe_pass: bool = True
|
||||
require_risk_pass: bool = True
|
||||
scale_in: list[ScaleStep] = Field(default_factory=list)
|
||||
profile_gate: ProfileGateConfig = Field(default_factory=ProfileGateConfig)
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _check(self) -> EntryConfig:
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
"""画像指标代码清单(闸门配置的唯一校验来源)。
|
||||
|
||||
放在 ``core`` 而不是 ``profile`` 里,是为了让 ``core.config``(策略配置校验)
|
||||
可以引用它,而 ``profile.pit`` 也能引用同一份定义 —— 避免出现
|
||||
「配置允许的指标」与「画像实际能算的指标」两个清单。
|
||||
|
||||
**为什么必须显式列清单**:闸门规则里写错一个指标名,若不在配置期拒绝,
|
||||
运行时就只会得到「该指标缺失 → 无法验证」,在保守策略下等于**永久不交易**。
|
||||
这类静默失效必须挡在配置校验里。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
#: 只需价格/每日指标即可计算的指标
|
||||
VALUATION_METRICS: frozenset[str] = frozenset(
|
||||
{"dv_yield", "ttm_dps", "close", "drawdown", "pe_ttm", "pb", "ps_ttm",
|
||||
"dv_vol_daily", "dv_vol_monthly", "dv_vol_quarterly", "dv_vol_annual"}
|
||||
)
|
||||
#: 需要风险/收益统计(仍只需价格 + 指数)
|
||||
RETURN_METRICS: frozenset[str] = frozenset(
|
||||
{"vol_250d", "max_drawdown_3y", "ret_1y", "ret_3y", "ret_5y", "cagr_5y", "beta_300"}
|
||||
)
|
||||
#: 需要分红明细
|
||||
DIVIDEND_METRICS: frozenset[str] = frozenset(
|
||||
{"dividend_continuity_years", "dividend_years_in_window", "dps",
|
||||
"dps_cagr_5y", "dps_volatility", "total_cash_dividend",
|
||||
"dividend_fy_free_cashflow"}
|
||||
)
|
||||
#: 需要财报(最重的查询:财务表各约 30 万行)
|
||||
FINANCIAL_METRICS: frozenset[str] = frozenset(
|
||||
{"roe", "roic", "gross_margin", "net_margin", "ocf_to_profit", "ocf_to_profit_calc",
|
||||
"roe_avg", "roic_avg", "gross_margin_avg", "net_margin_avg", "ocf_to_profit_avg",
|
||||
"debt_ratio", "free_cashflow", "n_income_attr_p", "total_assets",
|
||||
"payout_ratio", "fcf_dividend_cover"}
|
||||
)
|
||||
#: 需要流动性(20 日均成交额)
|
||||
LIQUIDITY_METRICS: frozenset[str] = frozenset({"avg_amount_20d"})
|
||||
|
||||
#: 有「历史分位」的指标 —— 即按窗口输出分布统计的那些。
|
||||
#: 其余是标量(只有最新值),对其做分位判断无意义,配置校验直接拒绝。
|
||||
PERCENTILE_METRICS: frozenset[str] = VALUATION_METRICS - {
|
||||
"dv_vol_daily", "dv_vol_monthly", "dv_vol_quarterly", "dv_vol_annual"
|
||||
}
|
||||
|
||||
#: 实时画像能产出的全部指标代码
|
||||
ALL_METRICS: frozenset[str] = (
|
||||
VALUATION_METRICS | RETURN_METRICS | DIVIDEND_METRICS
|
||||
| FINANCIAL_METRICS | LIQUIDITY_METRICS
|
||||
)
|
||||
|
||||
#: 观测单位是**交易日**的指标 —— 只有这些能用交易日历当覆盖率分母。
|
||||
#:
|
||||
#: 其余指标的 ``n_obs`` 不是交易日数:
|
||||
#: - 年报均值类(``roe_avg`` / ``ocf_to_profit_avg`` …)的 ``n_obs`` 是**财年数**(常为 3~5),
|
||||
#: 拿它比「5 年窗口应有 1219 个交易日」会算出 0.4% 这种荒谬覆盖率;
|
||||
#: - 标量快照类(``payout_ratio`` / ``debt_ratio`` …)的 ``n_obs`` 是 1。
|
||||
#:
|
||||
#: 量纲对不上就不能比 —— 这是本项目反复踩到的同一类错误。
|
||||
DAILY_OBSERVATION_METRICS: frozenset[str] = VALUATION_METRICS | RETURN_METRICS
|
||||
|
||||
#: 闸门规则允许引用的指标(与 ALL_METRICS 相同,单独命名以表达「对外契约」)
|
||||
GATE_METRICS: frozenset[str] = ALL_METRICS
|
||||
+99
-10
@@ -123,12 +123,16 @@ def _coverage_check(
|
||||
col: str,
|
||||
want_start: date,
|
||||
cfg: DataSourceConfig,
|
||||
min_symbols_per_day: int,
|
||||
min_symbols_per_day: int | None = None,
|
||||
sample_days: int = 6,
|
||||
) -> CheckResult:
|
||||
"""覆盖度检查。
|
||||
|
||||
性能说明:``stock_daily`` 有 770 万行,对全部交易日做
|
||||
``want_start`` 是「数据应至少覆盖到哪一天」的目标;``min_symbols_per_day``
|
||||
是**绝对下限的覆盖**(一般不用传,留空即按 ``datasource.yml: sync`` 的
|
||||
「按年份成比例」口径判定 —— 早年市场只有一千多只股票,固定阈值会误报)。
|
||||
|
||||
性能说明:``stock_daily`` 有 1,600 万行,对全部交易日做
|
||||
``GROUP BY trade_date`` + ``COUNT(DISTINCT symbol)`` 会触发全表聚合(数分钟)。
|
||||
因此改为**采样若干交易日**再取最小股票数 —— 足以发现「某段区间数据稀疏」的问题,
|
||||
成本从 O(全部行) 降到 O(采样日)。
|
||||
@@ -168,9 +172,15 @@ def _coverage_check(
|
||||
)
|
||||
counts.append((d, n))
|
||||
min_day, min_per_day = min(counts, key=lambda x: x[1])
|
||||
# 阈值必须**随年份变化**:2005 年 A 股只有约 1,350 只股票,
|
||||
# 用固定 2,000 只会把早年正常数据误报成「疑似数据稀疏」——
|
||||
# 与同步器断点续传曾经踩到的是同一个坑(见 sync.price.fetched_days)。
|
||||
floor, ratio, by_year = _expected_thresholds(cfg)
|
||||
need_min = max(floor, int(ratio * by_year.get(min_day.year, 0)))
|
||||
metrics["sampled_days"] = {str(d): n for d, n in counts}
|
||||
metrics["min_symbols_per_day"] = min_per_day
|
||||
metrics["min_symbols_date"] = str(min_day)
|
||||
metrics["min_symbols_expected"] = need_min
|
||||
metrics["total_trading_days"] = len(all_days)
|
||||
|
||||
if not enough_years:
|
||||
@@ -183,13 +193,13 @@ def _coverage_check(
|
||||
f"当前覆盖 {mn} ~ {mx}({len(all_days):,} 个交易日)",
|
||||
metrics,
|
||||
)
|
||||
if min_per_day < min_symbols_per_day:
|
||||
if min_per_day < need_min:
|
||||
return CheckResult(
|
||||
code,
|
||||
name,
|
||||
WARN,
|
||||
table,
|
||||
f"{min_day} 仅有 {min_per_day} 只股票(低于 {min_symbols_per_day}),疑似数据稀疏",
|
||||
f"{min_day} 仅有 {min_per_day} 只股票(按当年市场规模应 ≥ {need_min}),疑似数据稀疏",
|
||||
metrics,
|
||||
)
|
||||
return CheckResult(
|
||||
@@ -197,28 +207,38 @@ def _coverage_check(
|
||||
name,
|
||||
OK,
|
||||
table,
|
||||
f"覆盖 {mn} ~ {mx}({len(all_days):,} 个交易日),采样最少 {min_per_day} 只({min_day})",
|
||||
f"覆盖 {mn} ~ {mx}({len(all_days):,} 个交易日),"
|
||||
f"采样最少 {min_per_day} 只({min_day},当年应 ≥ {need_min})",
|
||||
metrics,
|
||||
)
|
||||
|
||||
|
||||
def _expected_thresholds(cfg: DataSourceConfig) -> tuple[int, float, dict[int, int]]:
|
||||
"""复用同步器的「按年份的规模阈值」,避免两处各写一套判定。"""
|
||||
from hdiv.data.sync.price import _day_thresholds
|
||||
|
||||
return _day_thresholds(cfg)
|
||||
|
||||
|
||||
def check_g2_price(cfg: DataSourceConfig) -> CheckResult:
|
||||
return _coverage_check(
|
||||
# 目标起点 = 2005:windows_years 最长 10 年,回测自 2015-01-05 起,
|
||||
# 要补满 10 年窗口就必须有 2005 年的数据
|
||||
"G2", "日线行情(含复权因子)", "stock_daily", "trade_date",
|
||||
date(2015, 1, 31), cfg, 2000,
|
||||
date(2005, 1, 31), cfg,
|
||||
)
|
||||
|
||||
|
||||
def check_g2b_adjust(cfg: DataSourceConfig) -> CheckResult:
|
||||
return _coverage_check(
|
||||
"G2b", "复权因子", "adjust_factor", "trade_date", date(2015, 1, 31), cfg, 2000
|
||||
"G2b", "复权因子", "adjust_factor", "trade_date", date(2005, 1, 31), cfg
|
||||
)
|
||||
|
||||
|
||||
def check_g3_daily_basic(cfg: DataSourceConfig) -> CheckResult:
|
||||
r = _coverage_check(
|
||||
"G3", "每日指标(PE/PB/股息率/市值)", "daily_basic", "trade_date",
|
||||
date(2015, 1, 31), cfg, 2000,
|
||||
date(2005, 1, 31), cfg,
|
||||
)
|
||||
if r.severity == OK:
|
||||
dv = _scalar(
|
||||
@@ -415,8 +435,8 @@ def check_g5b_index_weight(cfg: DataSourceConfig) -> CheckResult:
|
||||
def check_g6_trading(cfg: DataSourceConfig) -> list[CheckResult]:
|
||||
out: list[CheckResult] = []
|
||||
for code, table, label, want in [
|
||||
("G6", "hd_suspend", "停牌记录", date(2015, 1, 31)),
|
||||
("G6b", "hd_limit", "涨跌停价", date(2015, 1, 31)),
|
||||
("G6", "hd_suspend", "停牌记录", date(2010, 1, 31)),
|
||||
("G6b", "hd_limit", "涨跌停价", date(2010, 1, 31)),
|
||||
]:
|
||||
st = _table_stats(table, cfg)
|
||||
if not st.get("exists") or st.get("rows", 0) == 0:
|
||||
@@ -536,6 +556,74 @@ def check_units(cfg: DataSourceConfig) -> CheckResult:
|
||||
)
|
||||
|
||||
|
||||
def check_ohlcv_units(cfg: DataSourceConfig) -> CheckResult:
|
||||
"""``stock_daily`` 量价单位一致性(成交量/成交额)。
|
||||
|
||||
``stock_daily`` 是「追加进既有 qlib 库」的表:2015-01~2019 的行由本项目
|
||||
从 Tushare 回补(原始单位:手 / 千元),2020 起沿用 qlib 存量(股 / 元),
|
||||
2019 年同日混着两种。``min_avg_amount_20d`` 之类阈值是按「元」写的,
|
||||
所以这种混用会**静默**把早年流动性低估 1000 倍、把股票池清空 ——
|
||||
必须由审计主动发现(单位错误不会抛异常,只会让结果全错)。
|
||||
|
||||
判据:``成交额 / (成交量 × 收盘价)``,≈1 = 已换算,≈0.1 = 原始口径。
|
||||
取多个年份的样本日,任何一天存在原始口径行即判 WARN(不是 FAIL:
|
||||
读取层 :func:`hdiv.data.units.normalize_ohlcv_units` 已做兜底换算,
|
||||
但存量数据仍应择机修复,且新增写入不得再引入原始单位)。
|
||||
"""
|
||||
from hdiv.data.units import OHLCV_RAW, OHLCV_UNKNOWN, detect_ohlcv_units
|
||||
|
||||
df_max = db.read_sql("SELECT MAX(trade_date) AS d FROM stock_daily", cfg=cfg)
|
||||
if df_max.empty or df_max["d"].iloc[0] is None:
|
||||
return CheckResult("UNIT-OHLCV", "量价单位自检", WARN, "stock_daily", "stock_daily 无数据")
|
||||
last = pd.to_datetime(df_max["d"].iloc[0]).date()
|
||||
|
||||
# 每年取一个样本日:覆盖回补区间与 qlib 存量区间
|
||||
sample_days: list[date] = []
|
||||
for y in range(last.year, max(last.year - 12, 2009), -1):
|
||||
row = db.read_sql(
|
||||
"SELECT MAX(trade_date) AS d FROM stock_daily WHERE YEAR(trade_date) = :y",
|
||||
{"y": y}, cfg=cfg,
|
||||
)
|
||||
if not row.empty and row["d"].iloc[0] is not None:
|
||||
sample_days.append(pd.to_datetime(row["d"].iloc[0]).date())
|
||||
if not sample_days:
|
||||
return CheckResult("UNIT-OHLCV", "量价单位自检", WARN, "stock_daily", "无法取样")
|
||||
|
||||
per_day: list[dict[str, Any]] = []
|
||||
raw_days: list[str] = []
|
||||
for d in sample_days:
|
||||
s = db.read_sql(
|
||||
"SELECT symbol, close, volume, amount FROM stock_daily WHERE trade_date = :d",
|
||||
{"d": d}, cfg=cfg,
|
||||
)
|
||||
if s.empty:
|
||||
continue
|
||||
for c in ("close", "volume", "amount"):
|
||||
s[c] = pd.to_numeric(s[c], errors="coerce")
|
||||
unit = detect_ohlcv_units(s)
|
||||
n_raw = int((unit == OHLCV_RAW).sum())
|
||||
n_conv = int((unit != OHLCV_RAW).sum() - (unit == OHLCV_UNKNOWN).sum())
|
||||
per_day.append({"date": str(d), "rows": len(s), "raw": n_raw,
|
||||
"converted": n_conv,
|
||||
"unknown": int((unit == OHLCV_UNKNOWN).sum())})
|
||||
if n_raw:
|
||||
raw_days.append(f"{d}({n_raw}/{len(s)} 行为原始单位)")
|
||||
|
||||
detail = {"sampled_days": per_day}
|
||||
if raw_days:
|
||||
return CheckResult(
|
||||
"UNIT-OHLCV", "量价单位自检", WARN, "stock_daily",
|
||||
"存在 Tushare 原始单位(手/千元)的行:" + ";".join(raw_days[:6])
|
||||
+ "。读取层已兜底换算为「股/元」,但存量数据应择机重刷,"
|
||||
"且新增写入必须走 sync.price.daily_frame 的换算路径。",
|
||||
detail,
|
||||
)
|
||||
return CheckResult(
|
||||
"UNIT-OHLCV", "量价单位自检", OK, "stock_daily",
|
||||
f"{len(per_day)} 个抽样年份的量价单位一致(股/元)", detail,
|
||||
)
|
||||
|
||||
|
||||
def check_universe_ready(cfg: DataSourceConfig) -> CheckResult:
|
||||
"""第一版策略所需的最小数据条件是否齐备。"""
|
||||
needed = {
|
||||
@@ -575,6 +663,7 @@ ALL_CHECKS = [
|
||||
check_g5b_index_weight,
|
||||
check_g6_trading,
|
||||
check_units,
|
||||
check_ohlcv_units,
|
||||
check_duplicates,
|
||||
check_symbol_orphans,
|
||||
check_universe_ready,
|
||||
|
||||
+110
-29
@@ -29,6 +29,7 @@ from hdiv.data import db
|
||||
from hdiv.data.units import (
|
||||
normalize_financial_panel,
|
||||
normalize_market_panel,
|
||||
normalize_ohlcv_units,
|
||||
)
|
||||
|
||||
|
||||
@@ -47,6 +48,21 @@ class PanelSpec:
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _symbol_filter(
|
||||
symbols: list[str] | None, params: dict, col: str = "symbol"
|
||||
) -> str:
|
||||
"""构造 ``AND <col> IN (...)`` 片段并把值绑进 ``params``。
|
||||
|
||||
``symbols`` 为空(或 None)时返回空串 —— 与旧行为完全一致(全市场)。
|
||||
这是**纯性能开关**:不改变任何 PIT 语义,只把全表扫描缩到目标股票。
|
||||
"""
|
||||
if not symbols:
|
||||
return ""
|
||||
ph = ", ".join(f":sym{i}" for i in range(len(symbols)))
|
||||
params.update({f"sym{i}": s for i, s in enumerate(symbols)})
|
||||
return f" AND {col} IN ({ph})"
|
||||
|
||||
|
||||
def data_version(cfg: DataSourceConfig | None = None) -> str:
|
||||
"""数据版本指纹 = 关键表的最大日期摘要,写入每次 run 以支持复现。
|
||||
|
||||
@@ -231,6 +247,11 @@ class Repo:
|
||||
none —— 不复权原始价
|
||||
qfq —— 前复权(以区间末为基准)
|
||||
hfq —— 后复权(以区间首为基准)
|
||||
|
||||
成交量/成交额**一律归一化为「股 / 元」**再返回:``stock_daily`` 里
|
||||
2015-2019 的行是 Tushare 原始单位(手 / 千元),2020 起是「股 / 元」,
|
||||
直接使用会让按「元」写的流动性阈值低估 1000 倍。见
|
||||
:func:`hdiv.data.units.normalize_ohlcv_units`。
|
||||
"""
|
||||
if not symbols:
|
||||
return pd.DataFrame(columns=["symbol", "trade_date", "close", "open", "high", "low"])
|
||||
@@ -252,6 +273,7 @@ class Repo:
|
||||
df["trade_date"] = pd.to_datetime(df["trade_date"]).dt.date
|
||||
for c in ("open", "high", "low", "close", "volume", "amount", "factor"):
|
||||
df[c] = pd.to_numeric(df[c], errors="coerce")
|
||||
df, _diag = normalize_ohlcv_units(df)
|
||||
if adjust != "none":
|
||||
f = df["factor"].fillna(1.0)
|
||||
if adjust == "hfq":
|
||||
@@ -276,20 +298,42 @@ class Repo:
|
||||
df["trade_date"] = pd.to_datetime(df["trade_date"]).dt.date
|
||||
return df
|
||||
|
||||
def avg_amount(self, asof: date, window: int = 20) -> pd.DataFrame:
|
||||
"""asof 前 window 个交易日的日均成交额(流动性过滤)。"""
|
||||
def avg_amount(
|
||||
self, asof: date, window: int = 20, *, symbols: list[str] | None = None
|
||||
) -> pd.DataFrame:
|
||||
"""asof 前 window 个交易日的日均成交额(流动性过滤)。
|
||||
|
||||
返回的 ``avg_amount`` 单位是**元**。为此必须逐行归一化后再取均值 ——
|
||||
早期实现直接 ``AVG(amount)``,把 2015-2019 的「千元」与 2020 起的「元」
|
||||
混在一起平均,且与按「元」配置的阈值比较,低估 1000 倍。
|
||||
实测后果:``min_avg_amount_20d: 20000000`` 在 2015-2019 实际等价于
|
||||
「日均成交额 ≥ 200 亿元」,股票池被整体清空。
|
||||
|
||||
``symbols`` 为纯性能开关(全市场约 10 万行;限定几十只股票后降到毫秒级)。
|
||||
"""
|
||||
d0 = self.trading_day(asof)
|
||||
days = self.trading_days(d0 - timedelta(days=window * 3), d0)
|
||||
days = days[-window:]
|
||||
if not days:
|
||||
return pd.DataFrame(columns=["symbol", "avg_amount"])
|
||||
return pd.DataFrame(columns=["symbol", "avg_amount", "n"])
|
||||
params: dict = {"s": days[0], "e": days[-1]}
|
||||
sym_in = _symbol_filter(symbols, params)
|
||||
df = db.read_sql(
|
||||
"SELECT symbol, AVG(amount) AS avg_amount, COUNT(*) AS n "
|
||||
"FROM stock_daily WHERE trade_date BETWEEN :s AND :e GROUP BY symbol",
|
||||
{"s": days[0], "e": days[-1]},
|
||||
"SELECT symbol, close, volume, amount "
|
||||
"FROM stock_daily WHERE trade_date BETWEEN :s AND :e" + sym_in,
|
||||
params,
|
||||
cfg=self.cfg,
|
||||
)
|
||||
return df
|
||||
if df.empty:
|
||||
return pd.DataFrame(columns=["symbol", "avg_amount", "n"])
|
||||
for c in ("close", "volume", "amount"):
|
||||
df[c] = pd.to_numeric(df[c], errors="coerce")
|
||||
df, _diag = normalize_ohlcv_units(df)
|
||||
# n 沿用 COUNT(*) 语义(该窗口内有行情的交易日数),与归一化无关
|
||||
g = df.groupby("symbol", as_index=False).agg(
|
||||
avg_amount=("amount", "mean"), n=("amount", "size")
|
||||
)
|
||||
return g
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 停牌 / 涨跌停
|
||||
@@ -320,24 +364,31 @@ class Repo:
|
||||
# 财务数据(PIT)
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
def financial_panel(self, asof: date) -> pd.DataFrame:
|
||||
def financial_panel(
|
||||
self, asof: date, *, symbols: list[str] | None = None
|
||||
) -> pd.DataFrame:
|
||||
"""asof 时点可见的最新一期财务数据(``ann_date <= asof``)。
|
||||
|
||||
实现要点:用窗口函数取「报告期最新」的那条;
|
||||
``ann_date <= asof`` 保证不使用未公告数据。
|
||||
|
||||
``symbols`` 非空时只取这些股票 —— 纯粹的性能开关(财务表各约 30 万行,
|
||||
全表扫描约 1.7 秒;限定几十只股票后降到毫秒级)。**不改变 PIT 语义**:
|
||||
它只是在原有 WHERE 上追加一个 ``symbol IN (...)``,不会改变任何
|
||||
(symbol, end_date) 分区内的 ``ROW_NUMBER`` 结果。
|
||||
"""
|
||||
fin = self._latest_financial(asof, "hd_fina_indicator", ["roe", "roic", "debt_to_assets",
|
||||
"grossprofit_margin", "netprofit_margin",
|
||||
"ocf_to_profit"])
|
||||
"ocf_to_profit"], symbols=symbols)
|
||||
if fin.empty:
|
||||
return fin
|
||||
cf = self._latest_financial(asof, "hd_cashflow", ["n_cashflow_act", "free_cashflow",
|
||||
"c_pay_dist_dpcp_int_exp"])
|
||||
"c_pay_dist_dpcp_int_exp"], symbols=symbols)
|
||||
bs = self._latest_financial(asof, "hd_balancesheet", ["total_assets", "total_liab",
|
||||
"total_hldr_eqy_exc_min_int",
|
||||
"money_cap", "goodwill"])
|
||||
"money_cap", "goodwill"], symbols=symbols)
|
||||
inc = self._latest_financial(asof, "hd_income", ["total_revenue", "revenue", "n_income",
|
||||
"n_income_attr_p"])
|
||||
"n_income_attr_p"], symbols=symbols)
|
||||
out = fin
|
||||
for other in (cf, bs, inc):
|
||||
if other.empty:
|
||||
@@ -352,7 +403,10 @@ class Repo:
|
||||
out = self._derive_financial(out)
|
||||
return out
|
||||
|
||||
def _latest_financial(self, asof: date, table: str, cols: list[str]) -> pd.DataFrame:
|
||||
def _latest_financial(
|
||||
self, asof: date, table: str, cols: list[str],
|
||||
*, symbols: list[str] | None = None,
|
||||
) -> pd.DataFrame:
|
||||
if not db.table_exists(table, self.cfg):
|
||||
return pd.DataFrame(columns=["symbol", "end_date", "ann_date", *cols])
|
||||
rt = (
|
||||
@@ -361,6 +415,8 @@ class Repo:
|
||||
else ""
|
||||
)
|
||||
sel = ", ".join(cols)
|
||||
params: dict = {"asof": asof}
|
||||
sym_cond = _symbol_filter(symbols, params)
|
||||
sql = f"""
|
||||
SELECT symbol, end_date, ann_date, {sel} FROM (
|
||||
SELECT symbol, end_date, ann_date, {sel},
|
||||
@@ -368,13 +424,13 @@ class Repo:
|
||||
PARTITION BY symbol ORDER BY end_date DESC, ann_date DESC
|
||||
) AS rn
|
||||
FROM {table}
|
||||
WHERE ann_date <= :asof {rt}
|
||||
WHERE ann_date <= :asof {rt} {sym_cond}
|
||||
-- 公告日不可能早于报告期;这类行是数据源错误(实测 920185.BJ 有 2 条),
|
||||
-- 若不过滤会构成未来函数。审计中仍会如实报告其数量。
|
||||
AND ann_date >= end_date
|
||||
) t WHERE rn = 1
|
||||
"""
|
||||
df = db.read_sql(sql, {"asof": asof}, cfg=self.cfg)
|
||||
df = db.read_sql(sql, params, cfg=self.cfg)
|
||||
for c in ("end_date", "ann_date"):
|
||||
if c in df.columns and not df.empty:
|
||||
df[c] = pd.to_datetime(df[c]).dt.date
|
||||
@@ -449,7 +505,9 @@ class Repo:
|
||||
df[c] = pd.to_numeric(df[c], errors="coerce")
|
||||
return df
|
||||
|
||||
def annual_financial_history(self, asof: date, *, years: int = 6) -> pd.DataFrame:
|
||||
def annual_financial_history(
|
||||
self, asof: date, *, years: int = 6, symbols: list[str] | None = None
|
||||
) -> pd.DataFrame:
|
||||
"""asof 时点可见的**年度**财务指标历史(用于 5 年平均等长期口径)。
|
||||
|
||||
为什么必须用年报而不是最新季报:
|
||||
@@ -458,22 +516,29 @@ class Repo:
|
||||
会把几乎所有好公司误杀(实测:600036.SH 的 2024Q1 ROE 仅 3.47%)。
|
||||
|
||||
因此这里只取 ``end_date`` 为 12-31 的年报,且 ``ann_date <= asof``(PIT)。
|
||||
|
||||
``symbols`` 是纯性能开关(见 :func:`_symbol_filter`):内层子查询与外层
|
||||
用同一个 symbol 过滤条件,因此不会改变任何 ``(symbol, end_date)`` 分组
|
||||
的 ``MAX(ann_date)``,结果与全市场口径逐行一致。
|
||||
"""
|
||||
if not db.table_exists("hd_fina_indicator", self.cfg):
|
||||
return pd.DataFrame(columns=["symbol", "year", "roe", "roic"])
|
||||
|
||||
since = date(asof.year - years - 1, 12, 31)
|
||||
params: dict = {"asof": asof, "since": since}
|
||||
sym_in = _symbol_filter(symbols, params)
|
||||
sym_in_f = _symbol_filter(symbols, params, col="f.symbol")
|
||||
fin = db.read_sql(
|
||||
"SELECT f.symbol, f.end_date, f.ann_date, f.roe, f.roic, f.grossprofit_margin, "
|
||||
" f.netprofit_margin, f.ocf_to_profit "
|
||||
"FROM hd_fina_indicator f "
|
||||
"JOIN (SELECT symbol, end_date, MAX(ann_date) AS a FROM hd_fina_indicator "
|
||||
" WHERE ann_date <= :asof AND MONTH(end_date) = 12 AND end_date >= :since "
|
||||
" AND ann_date >= end_date "
|
||||
" AND ann_date >= end_date" + sym_in + " "
|
||||
" GROUP BY symbol, end_date) m "
|
||||
" ON m.symbol = f.symbol AND m.end_date = f.end_date AND m.a = f.ann_date "
|
||||
"WHERE MONTH(f.end_date) = 12 AND f.end_date >= :since",
|
||||
{"asof": asof, "since": since},
|
||||
"WHERE MONTH(f.end_date) = 12 AND f.end_date >= :since" + sym_in_f,
|
||||
params,
|
||||
cfg=self.cfg,
|
||||
)
|
||||
if fin.empty:
|
||||
@@ -484,23 +549,27 @@ class Repo:
|
||||
fin[c] = pd.to_numeric(fin[c], errors="coerce") / 100.0
|
||||
|
||||
# 经营现金流/净利润:用现金流量表与利润表年报口径补算(比 fina_indicator 更可靠)
|
||||
ocf = self._annual_ocf_ratio(asof, since)
|
||||
ocf = self._annual_ocf_ratio(asof, since, symbols=symbols)
|
||||
if not ocf.empty:
|
||||
fin = fin.merge(ocf, on=["symbol", "year"], how="left")
|
||||
return fin
|
||||
|
||||
def _annual_ocf_ratio(self, asof: date, since: date) -> pd.DataFrame:
|
||||
def _annual_ocf_ratio(
|
||||
self, asof: date, since: date, *, symbols: list[str] | None = None
|
||||
) -> pd.DataFrame:
|
||||
"""年报口径的 经营现金流 / 归母净利润。"""
|
||||
if not (db.table_exists("hd_cashflow", self.cfg) and db.table_exists("hd_income", self.cfg)):
|
||||
return pd.DataFrame(columns=["symbol", "year", "ocf_to_netprofit_calc"])
|
||||
params: dict = {"asof": asof, "since": since}
|
||||
sym_in = _symbol_filter(symbols, params, col="c.symbol")
|
||||
df = db.read_sql(
|
||||
"SELECT c.symbol, c.end_date, c.n_cashflow_act, i.n_income_attr_p "
|
||||
"FROM hd_cashflow c "
|
||||
"JOIN hd_income i ON i.symbol = c.symbol AND i.end_date = c.end_date "
|
||||
" AND i.ann_date = c.ann_date AND i.report_type = c.report_type "
|
||||
"WHERE c.report_type = '1' AND MONTH(c.end_date) = 12 "
|
||||
" AND c.end_date >= :since AND c.ann_date <= :asof",
|
||||
{"asof": asof, "since": since},
|
||||
" AND c.end_date >= :since AND c.ann_date <= :asof" + sym_in,
|
||||
params,
|
||||
cfg=self.cfg,
|
||||
)
|
||||
if df.empty:
|
||||
@@ -512,7 +581,8 @@ class Repo:
|
||||
return df[["symbol", "year", "ocf_to_netprofit_calc"]]
|
||||
|
||||
def annual_financial_averages(
|
||||
self, asof: date, *, years: int = 5, min_years: int = 3
|
||||
self, asof: date, *, years: int = 5, min_years: int = 3,
|
||||
symbols: list[str] | None = None, hist: pd.DataFrame | None = None,
|
||||
) -> pd.DataFrame:
|
||||
"""把年报历史聚合成「N 年平均」指标。
|
||||
|
||||
@@ -520,8 +590,14 @@ class Repo:
|
||||
``net_margin_avg`` / ``ocf_to_profit_avg`` / ``fin_years_count`` /
|
||||
``fin_latest_year``。样本年数不足 ``min_years`` 时对应平均值为 NaN
|
||||
(宁可标注「不可得」,也不要用不足的样本猜)。
|
||||
|
||||
``hist`` 允许调用方传入已取好的 :meth:`annual_financial_history` 结果,
|
||||
避免同一时点重复查询(PIT 画像逐日调用时这是 2 倍开销)。
|
||||
"""
|
||||
hist = self.annual_financial_history(asof, years=years)
|
||||
hist = (
|
||||
self.annual_financial_history(asof, years=years, symbols=symbols)
|
||||
if hist is None else hist
|
||||
)
|
||||
if hist.empty:
|
||||
return pd.DataFrame(
|
||||
columns=["symbol", "roe_avg", "roic_avg", "gross_margin_avg",
|
||||
@@ -552,7 +628,9 @@ class Repo:
|
||||
out.loc[thin, c] = pd.NA
|
||||
return out
|
||||
|
||||
def annual_financials(self, asof: date, *, years: int = 12) -> pd.DataFrame:
|
||||
def annual_financials(
|
||||
self, asof: date, *, years: int = 12, symbols: list[str] | None = None
|
||||
) -> pd.DataFrame:
|
||||
"""按**财年**对齐的年度财务数据(PIT)。
|
||||
|
||||
为什么必须单独提供:``financial_panel`` 返回的是**最新一期**财报
|
||||
@@ -562,7 +640,7 @@ class Repo:
|
||||
|
||||
返回列:``symbol / year / n_income_attr_p / n_cashflow_act /
|
||||
free_cashflow / c_pay_dist_dpcp_int_exp``,仅取年报(``MONTH(end_date)=12``),
|
||||
且 ``ann_date <= asof``。
|
||||
且 ``ann_date <= asof``。``symbols`` 为纯性能开关。
|
||||
"""
|
||||
if not (db.table_exists("hd_income", self.cfg) and db.table_exists("hd_cashflow", self.cfg)):
|
||||
return pd.DataFrame(
|
||||
@@ -570,6 +648,8 @@ class Repo:
|
||||
"free_cashflow", "c_pay_dist_dpcp_int_exp"]
|
||||
)
|
||||
since = date(asof.year - years - 1, 12, 31)
|
||||
params: dict = {"asof": asof, "since": since}
|
||||
sym_in = _symbol_filter(symbols, params, col="i.symbol")
|
||||
df = db.read_sql(
|
||||
"SELECT i.symbol, i.end_date, i.n_income, i.n_income_attr_p, "
|
||||
" c.n_cashflow_act, c.free_cashflow, c.c_pay_dist_dpcp_int_exp "
|
||||
@@ -577,8 +657,9 @@ class Repo:
|
||||
"LEFT JOIN hd_cashflow c ON c.symbol = i.symbol AND c.end_date = i.end_date "
|
||||
" AND c.ann_date = i.ann_date AND c.report_type = i.report_type "
|
||||
"WHERE i.report_type = '1' AND MONTH(i.end_date) = 12 "
|
||||
" AND i.end_date >= :since AND i.ann_date <= :asof AND i.ann_date >= i.end_date",
|
||||
{"asof": asof, "since": since},
|
||||
" AND i.end_date >= :since AND i.ann_date <= :asof AND i.ann_date >= i.end_date"
|
||||
+ sym_in,
|
||||
params,
|
||||
cfg=self.cfg,
|
||||
)
|
||||
if df.empty:
|
||||
|
||||
+73
-10
@@ -21,6 +21,7 @@ from hdiv.core.config import DataSourceConfig, load_config
|
||||
from hdiv.data import db
|
||||
from hdiv.data.sync.base import sync_job, to_date, to_float, upsert
|
||||
from hdiv.data.tushare_client import TushareClient
|
||||
from hdiv.data.units import amount_qian_to_yuan, vol_shou_to_shares
|
||||
|
||||
EXCHANGES = ("SSE", "SZSE", "BSE")
|
||||
EXCHANGE_CODE = {"SSE": "SH", "SZSE": "SZ", "BSE": "BJ"}
|
||||
@@ -60,19 +61,59 @@ def open_days(start: date, end: date, cfg: DataSourceConfig | None = None) -> li
|
||||
|
||||
|
||||
#: 一个交易日被视为「已同步完成」所需的最少股票数。
|
||||
#: A 股自 2015 年起每个交易日都有 2,000 只以上在交易。
|
||||
#: 保留为**绝对下限**:A 股自 2015 年起每个交易日都有 2,000 只以上在交易。
|
||||
#: 早期实现只看「该日期是否存在」,会把**只填了几百只**的半成品日当成已完成
|
||||
#: (实测:原 qlib 数据在 2019 年仅 243 只/日,被误判为已同步,形成整年数据空洞)。
|
||||
#: 但现在**不再单独使用它** —— 2005 年 A 股只有约 1,350 只,固定 1,500 会让
|
||||
#: 2005-2009 的每一天都判为「未完成」,断点续传失效。实际阈值见
|
||||
#: :func:`_day_threshold`:``max(绝对下限, 比例 × 当年应有上市股票数)``。
|
||||
MIN_SYMBOLS_PER_DAY = 1500
|
||||
|
||||
|
||||
def _expected_symbols_by_year(cfg: DataSourceConfig) -> dict[int, int]:
|
||||
"""``{年份: 该年末累计上市股票数}``(来自 ``stock.list_date``,独立于行情表)。
|
||||
|
||||
用独立的 ``stock`` 表当参照,而不是行情表自身的观测数 —— 后者是循环论证:
|
||||
整段缺失的年份根本没有行,无从判断它「应该有」多少。
|
||||
忽略退市会让这个数偏大(是上界),0.6 的比例留了足够余量。
|
||||
"""
|
||||
try:
|
||||
df = db.read_sql(
|
||||
"SELECT YEAR(list_date) AS y, COUNT(*) AS n FROM stock "
|
||||
"WHERE list_date IS NOT NULL AND YEAR(list_date) > 1990 GROUP BY y",
|
||||
cfg=cfg,
|
||||
)
|
||||
except Exception:
|
||||
return {}
|
||||
if df.empty:
|
||||
return {}
|
||||
counts = {int(r["y"]): int(r["n"]) for _, r in df.iterrows()}
|
||||
cum, out = 0, {}
|
||||
for y in range(min(counts), max(counts) + 1):
|
||||
cum += counts.get(y, 0)
|
||||
out[y] = cum
|
||||
return out
|
||||
|
||||
|
||||
def _day_thresholds(cfg: DataSourceConfig) -> tuple[int, float, dict[int, int]]:
|
||||
scfg = getattr(cfg, "sync", None)
|
||||
floor = int(getattr(scfg, "min_symbols_floor", 200))
|
||||
ratio = float(getattr(scfg, "min_symbols_ratio", 0.6))
|
||||
return floor, ratio, _expected_symbols_by_year(cfg)
|
||||
|
||||
|
||||
def fetched_days(
|
||||
table: str, start: date, end: date, cfg: DataSourceConfig, *, min_symbols: int = MIN_SYMBOLS_PER_DAY
|
||||
table: str, start: date, end: date, cfg: DataSourceConfig,
|
||||
*, min_symbols: int | None = None,
|
||||
) -> set[date]:
|
||||
"""**数据完整**的交易日集合(用于断点续传)。
|
||||
|
||||
判定标准是「当日股票数 >= min_symbols」,而不是「当日是否存在行」——
|
||||
判定标准是「当日股票数 >= 该日应有的规模」,而不是「当日是否存在行」——
|
||||
否则半成品日期会被跳过,留下难以察觉的数据空洞。
|
||||
|
||||
阈值随年份变化(见 ``SyncConfig``):早年 A 股只有一千多只股票,
|
||||
固定阈值会把 2005-2009 的每一天都判成「未完成」,断点续传失效。
|
||||
``min_symbols`` 显式给定时按旧口径(固定阈值)判定,保持向后兼容。
|
||||
"""
|
||||
df = db.read_sql(
|
||||
f"SELECT trade_date AS d, COUNT(DISTINCT symbol) AS n FROM `{table}` "
|
||||
@@ -82,11 +123,19 @@ def fetched_days(
|
||||
)
|
||||
if df.empty:
|
||||
return set()
|
||||
return {
|
||||
to_date(r["d"]) # type: ignore[misc]
|
||||
for _, r in df.iterrows()
|
||||
if int(r["n"]) >= min_symbols
|
||||
}
|
||||
floor, ratio, by_year = _day_thresholds(cfg)
|
||||
out: set[date] = set()
|
||||
for _, r in df.iterrows():
|
||||
d = to_date(r["d"])
|
||||
if d is None:
|
||||
continue
|
||||
if min_symbols is not None:
|
||||
need = int(min_symbols)
|
||||
else:
|
||||
need = max(floor, int(ratio * by_year.get(d.year, 0)))
|
||||
if int(r["n"]) >= need:
|
||||
out.add(d)
|
||||
return out
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -121,6 +170,14 @@ def _tp(rows: list[dict]) -> tuple[list[date], list[str]]:
|
||||
|
||||
|
||||
def daily_frame(rows: list[dict]) -> pd.DataFrame:
|
||||
"""Tushare ``daily`` → ``stock_daily`` 行,**并统一量价单位**。
|
||||
|
||||
Tushare 的 ``vol`` 是「手」、``amount`` 是「千元」,而既有的 ``stock_daily``
|
||||
(qlib 存量数据)是「股 / 元」。早期实现原样写入,于是同一列在 2015-2019 与
|
||||
2020 起是两套单位,按「元」写的流动性阈值在早年低估 1000 倍。
|
||||
这里在写入端就换算,读取端的 :func:`hdiv.data.units.normalize_ohlcv_units`
|
||||
作为存量数据的兜底(对已换算行幂等)。
|
||||
"""
|
||||
if not rows:
|
||||
return pd.DataFrame()
|
||||
td, sym = _tp(rows)
|
||||
@@ -133,8 +190,14 @@ def daily_frame(rows: list[dict]) -> pd.DataFrame:
|
||||
"high": [to_float(r.get("high")) for r in rows],
|
||||
"low": [to_float(r.get("low")) for r in rows],
|
||||
"close": [to_float(r.get("close")) for r in rows],
|
||||
"volume": [to_float(r.get("vol")) for r in rows],
|
||||
"amount": [to_float(r.get("amount")) for r in rows],
|
||||
# 手 → 股
|
||||
"volume": vol_shou_to_shares(
|
||||
pd.Series([to_float(r.get("vol")) for r in rows], dtype="float64")
|
||||
),
|
||||
# 千元 → 元
|
||||
"amount": amount_qian_to_yuan(
|
||||
pd.Series([to_float(r.get("amount")) for r in rows], dtype="float64")
|
||||
),
|
||||
"source": "tushare",
|
||||
"adjust": "none",
|
||||
}
|
||||
|
||||
@@ -81,7 +81,14 @@ def sync_trading_constraints(
|
||||
cfg = load_config("datasource")
|
||||
days = open_days(start, end, cfg)
|
||||
if resume:
|
||||
done_s = fetched_days(SUSPEND_TABLE, start, end, cfg)
|
||||
# 两张表的「完整性」判定不能用同一把尺子:
|
||||
# hd_limit —— 每个交易日有全市场约 2,600~3,500 行,可用按年份的规模阈值;
|
||||
# hd_suspend —— 每天只有**当天停牌的那几十只**(实测 18~260 行),
|
||||
# 永远达不到「市场规模的 60%」,于是每一天都被判成「未完成」,
|
||||
# 断点续传彻底失效(每次重跑都重新拉全部停牌日)。
|
||||
# 停牌表只能退化为「有行即视为已同步」;真正无停牌的交易日会被重复拉取,
|
||||
# 代价极小(每天 1 次调用且返回空)。
|
||||
done_s = fetched_days(SUSPEND_TABLE, start, end, cfg, min_symbols=1)
|
||||
done_l = fetched_days(LIMIT_TABLE, start, end, cfg)
|
||||
days_s = [d for d in days if d not in done_s]
|
||||
days_l = [d for d in days if d not in done_l]
|
||||
|
||||
@@ -66,6 +66,111 @@ def vol_shou_to_shares(s: pd.Series) -> pd.Series:
|
||||
return pd.to_numeric(s, errors="coerce") * SHOU
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# stock_daily 的量价单位(历史遗留:同一列混着两种单位)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
#: ``stock_daily`` 中一行已经处于本项目统一口径(成交量=股,成交额=元)。
|
||||
OHLCV_CONVERTED = "converted"
|
||||
#: ``stock_daily`` 中一行仍是 Tushare 原始口径(成交量=手,成交额=千元)。
|
||||
OHLCV_RAW = "raw"
|
||||
#: 无法判定(缺 close/volume/amount,或非正值)。
|
||||
OHLCV_UNKNOWN = "unknown"
|
||||
|
||||
#: raw 与 converted 的比值相差正好 10 倍(见 :func:`ohlcv_unit_ratio`),
|
||||
#: 而 A 股有 ±10% 涨跌幅限制使 VWAP/收盘价落在 0.9~1.1,所以 0.3 这个阈值
|
||||
#: 有 ~3 倍的安全边际 —— 不会把正常行情误判成另一种单位。
|
||||
_RAW_MAX_RATIO = 0.3
|
||||
|
||||
|
||||
def ohlcv_unit_ratio(df: pd.DataFrame) -> pd.Series:
|
||||
"""``成交额 / (成交量 × 收盘价)``:≈1 表示已换算,≈0.1 表示 Tushare 原始单位。
|
||||
|
||||
这是唯一可靠的判据:列名完全看不出单位,而量级会。推导:
|
||||
|
||||
- 统一口径(股 / 元):``amount / (volume × close) = 1``
|
||||
- 原始口径(手 / 千元):``amount / (volume × close) = 100 / 1000 = 0.1``
|
||||
|
||||
``VWAP = amount / volume`` 与 ``close`` 的比值受涨跌幅限制约束,
|
||||
因此该比值只有 1 或 0.1 两个可能,不存在中间态。
|
||||
"""
|
||||
need = {"amount", "volume", "close"}
|
||||
if df.empty or not need <= set(df.columns):
|
||||
return pd.Series(dtype="float64", index=df.index)
|
||||
amt = pd.to_numeric(df["amount"], errors="coerce")
|
||||
vol = pd.to_numeric(df["volume"], errors="coerce")
|
||||
close = pd.to_numeric(df["close"], errors="coerce")
|
||||
denom = vol * close
|
||||
ok = (denom > 0) & amt.notna()
|
||||
out = pd.Series(float("nan"), index=df.index, dtype="float64")
|
||||
out[ok] = amt[ok] / denom[ok]
|
||||
return out
|
||||
|
||||
|
||||
def detect_ohlcv_units(df: pd.DataFrame) -> pd.Series:
|
||||
"""逐行判定 ``stock_daily`` 的量价单位,返回 ``raw`` / ``converted`` / ``unknown``。"""
|
||||
if df.empty:
|
||||
return pd.Series(dtype="object", index=df.index)
|
||||
r = ohlcv_unit_ratio(df)
|
||||
out = pd.Series(OHLCV_UNKNOWN, index=df.index, dtype="object")
|
||||
out[r.notna() & (r > _RAW_MAX_RATIO)] = OHLCV_CONVERTED
|
||||
out[r.notna() & (r <= _RAW_MAX_RATIO)] = OHLCV_RAW
|
||||
return out
|
||||
|
||||
|
||||
def normalize_ohlcv_units(
|
||||
df: pd.DataFrame,
|
||||
*,
|
||||
volume_col: str = "volume",
|
||||
amount_col: str = "amount",
|
||||
close_col: str = "close",
|
||||
) -> tuple[pd.DataFrame, dict[str, Any]]:
|
||||
"""把 ``stock_daily`` 的成交量/成交额统一到「股 / 元」。
|
||||
|
||||
**为什么必须在读取时做**:``stock_daily`` 是「追加进既有 qlib 库」的表 ——
|
||||
2015-01~2019 的行由本项目从 Tushare 回补,写的是原始单位(手 / 千元);
|
||||
2020 起的行沿用 qlib 既有数据(股 / 元);2019 年同日混着两种。
|
||||
而 ``min_avg_amount_20d`` 这类阈值是按「元」写的,于是 2015-2019 的
|
||||
20 日均额被低估 1000 倍 —— 流动性门槛实际变成「日均成交额 ≥ 200 亿元」,
|
||||
把 2015-2019 的股票池整体清空(实测 2016/2017/2018 各筛选出 0 只)。
|
||||
|
||||
该函数是**幂等**的:已换算的行比值 ≈1,不会被二次换算。
|
||||
返回 ``(新 DataFrame, 诊断信息)``,不修改入参。
|
||||
"""
|
||||
if df.empty:
|
||||
return df, {"total": 0, "raw": 0, "converted": 0, "unknown": 0, "fixed": 0}
|
||||
for col in (volume_col, amount_col, close_col):
|
||||
if col not in df.columns:
|
||||
# 缺少任一列都无法判定单位,只能原样返回(并在诊断里体现)
|
||||
return df, {
|
||||
"total": int(len(df)), "raw": 0, "converted": 0,
|
||||
"unknown": int(len(df)), "fixed": 0,
|
||||
"error": f"缺少列 {col},无法判定单位",
|
||||
}
|
||||
unit = detect_ohlcv_units(
|
||||
df.rename(columns={volume_col: "volume", amount_col: "amount",
|
||||
close_col: "close"})
|
||||
)
|
||||
out = df.copy()
|
||||
raw_mask = unit.to_numpy() == OHLCV_RAW
|
||||
if raw_mask.any():
|
||||
out.loc[raw_mask, volume_col] = (
|
||||
pd.to_numeric(out.loc[raw_mask, volume_col], errors="coerce") * SHOU
|
||||
)
|
||||
out.loc[raw_mask, amount_col] = (
|
||||
pd.to_numeric(out.loc[raw_mask, amount_col], errors="coerce") * QIAN
|
||||
)
|
||||
counts = unit.value_counts()
|
||||
diag = {
|
||||
"total": int(len(df)),
|
||||
"raw": int(counts.get(OHLCV_RAW, 0)),
|
||||
"converted": int(counts.get(OHLCV_CONVERTED, 0)),
|
||||
"unknown": int(counts.get(OHLCV_UNKNOWN, 0)),
|
||||
"fixed": int(raw_mask.sum()),
|
||||
}
|
||||
return out, diag
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 面板归一化
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
+89
-12
@@ -37,6 +37,11 @@ from hdiv.factor.dividend_yield import (
|
||||
rolling_volatility,
|
||||
window_slice,
|
||||
)
|
||||
from hdiv.profile.coverage import (
|
||||
expected_by_window,
|
||||
format_warning,
|
||||
summarise,
|
||||
)
|
||||
|
||||
# 指标展示名与单位(呈现层用)
|
||||
METRIC_META: dict[str, dict[str, str]] = {
|
||||
@@ -55,6 +60,13 @@ METRIC_META: dict[str, dict[str, str]] = {
|
||||
"dps": {"label": "每股分红", "unit": "price"},
|
||||
"payout_ratio": {"label": "分红支付率", "unit": "pct"},
|
||||
"fcf_dividend_cover": {"label": "FCF 对分红覆盖", "unit": "ratio"},
|
||||
# 两个「自由现金流」是不同的口径,必须分开:
|
||||
# free_cashflow —— 最近一期已公告财报的自由现金流
|
||||
# dividend_fy_free_cashflow —— 与最近一次分红**同一财年**的自由现金流
|
||||
# 后者才是计算 FCF 覆盖倍数的分子;混用一个代码会让同一指标在不同报告里
|
||||
# 显示两个不同的数(实测格力电器 2018-05-18:67.1 亿 vs 70.5 亿)。
|
||||
"free_cashflow": {"label": "自由现金流(最近一期)", "unit": "money"},
|
||||
"dividend_fy_free_cashflow": {"label": "自由现金流(分红同财年)", "unit": "money"},
|
||||
"dividend_continuity_years": {"label": "连续分红年数", "unit": "years"},
|
||||
"dps_cagr_5y": {"label": "DPS 5年复合增速", "unit": "pct"},
|
||||
"dps_volatility": {"label": "DPS 波动率", "unit": "ratio"},
|
||||
@@ -102,7 +114,15 @@ class ProfileBuilder:
|
||||
asof: str | date | None = None,
|
||||
persist: bool = True,
|
||||
verbose: bool = True,
|
||||
return_rows: bool = False,
|
||||
) -> dict[str, Any]:
|
||||
"""计算画像。
|
||||
|
||||
``return_rows=True`` 时在结果里附带 ``stat_rows`` / ``series_rows`` /
|
||||
``score_rows`` 原始行(不落库也能检查逐指标取值)。这是给**等价性测试**
|
||||
用的:实时画像(``profile.pit``)与批量画像必须逐值一致,而验证这一点
|
||||
需要一个不写库也能读到逐指标结果的入口。
|
||||
"""
|
||||
cfg = load_config("datasource")
|
||||
c = self.config
|
||||
|
||||
@@ -202,6 +222,20 @@ class ProfileBuilder:
|
||||
if verbose and i % 50 == 0:
|
||||
print(f" [{i}/{len(syms)}] 已画像", flush=True)
|
||||
|
||||
# 窗口覆盖率自检:名义窗口与实际可用数据是两件事。
|
||||
# 数据起点晚于窗口左端时窗口会被**静默截短**(实测 2018-05-18 的
|
||||
# 「5 年」窗口只有 3.4 年),而 n_obs 门槛低到 20 就放行 ——
|
||||
# 这里至少把它说出来,避免把「3.4 年」当成「5 年」用。
|
||||
cov_summary = summarise(
|
||||
stat_rows,
|
||||
expected_by_window(
|
||||
self.repo, asof_d,
|
||||
# 除配置窗口外,收益/回撤类指标还会产出 1/3 年窗口
|
||||
sorted({int(w) for w in c.windows_years} | {1, 3}),
|
||||
),
|
||||
)
|
||||
cov_warn = format_warning(cov_summary)
|
||||
|
||||
result = {
|
||||
"run_id": run_id,
|
||||
"asof_date": asof_d,
|
||||
@@ -211,7 +245,15 @@ class ProfileBuilder:
|
||||
"score_count": len(score_rows),
|
||||
"skipped": skipped,
|
||||
"symbols": sorted({r["symbol"] for r in stat_rows}),
|
||||
"window_coverage": cov_summary,
|
||||
}
|
||||
if cov_warn:
|
||||
result["warnings"] = [cov_warn]
|
||||
if return_rows:
|
||||
# 不落库也能逐指标核对(等价性测试用)
|
||||
result["stat_rows"] = stat_rows
|
||||
result["series_rows"] = series_rows
|
||||
result["score_rows"] = score_rows
|
||||
|
||||
if persist:
|
||||
result["written"] = self._persist(
|
||||
@@ -244,11 +286,25 @@ class ProfileBuilder:
|
||||
fin_latest: pd.DataFrame,
|
||||
div_by_symbol: dict[str, list[dict]] | None = None,
|
||||
fy_table: dict[tuple[str, int], Any] | None = None,
|
||||
build_series: bool = True,
|
||||
) -> dict[str, Any] | None:
|
||||
"""单股画像。
|
||||
|
||||
``build_series=False`` 跳过**仅供绘图**的降采样序列,统计量完全不受影响。
|
||||
实时画像(``profile.pit``)只需要 ``current_value`` / ``current_percentile``,
|
||||
不需要画图 —— 而构造那几条最多 1500 点的序列占掉本函数约一半耗时
|
||||
(实测每个快照 0.86s → 0.37s)。
|
||||
|
||||
**本函数强制按 ``trade_date`` 排序**,不依赖调用方给有序帧:所有序列的
|
||||
「当日值」都是取最后一个观测(``current = sub.iloc[-1]``),行序错了就会
|
||||
取到任意一天的值,而且**不会报错**。2026-10-04 回补数据时就踩到过
|
||||
(详见 :meth:`_load_daily_basic`)。
|
||||
"""
|
||||
c = self.config
|
||||
px = price[price["symbol"] == sym]
|
||||
if px.empty:
|
||||
return None
|
||||
px = px.sort_values("trade_date") # 见 docstring:行序即语义
|
||||
close = px.set_index("trade_date")["close"]
|
||||
close.index = pd.to_datetime(close.index)
|
||||
|
||||
@@ -258,6 +314,7 @@ class ProfileBuilder:
|
||||
events.get(sym, pd.DataFrame()),
|
||||
ttm_days=c.ttm_dividend.window_days,
|
||||
grace_days=c.ttm_dividend.grace_days,
|
||||
smooth_spikes=c.ttm_dividend.smooth_spikes,
|
||||
)
|
||||
if yser.empty:
|
||||
return None
|
||||
@@ -271,6 +328,7 @@ class ProfileBuilder:
|
||||
}
|
||||
bs = basics[basics["symbol"] == sym]
|
||||
if not bs.empty:
|
||||
bs = bs.sort_values("trade_date") # 同上:行序即「当日值」的语义
|
||||
b = bs.set_index(pd.to_datetime(bs["trade_date"]))
|
||||
for col in ("pe_ttm", "pb", "ps_ttm"):
|
||||
if col in b.columns:
|
||||
@@ -309,8 +367,8 @@ class ProfileBuilder:
|
||||
"OK" if st.get("n_obs", 0) >= min(c.min_obs_days, 20) else "INSUFFICIENT",
|
||||
)
|
||||
)
|
||||
# 展示序列只落配置指定的指标
|
||||
if metric in c.series_metrics:
|
||||
# 展示序列只落配置指定的指标(且仅在需要时构造 —— 见 build_series)
|
||||
if build_series and metric in c.series_metrics:
|
||||
disp = resample_for_storage(
|
||||
pd.DataFrame({"trade_date": s.index, "value": s.to_numpy()}),
|
||||
c.series_max_points,
|
||||
@@ -452,17 +510,25 @@ class ProfileBuilder:
|
||||
st.update(DividendFilter._payout_and_cover(recs, fy_row, asof, c))
|
||||
|
||||
out: list[dict[str, Any]] = []
|
||||
# 标量型质量指标
|
||||
for code in (
|
||||
"dividend_continuity_years", "dividend_years_in_window",
|
||||
"ttm_dps", "dps_cagr_5y", "dps_volatility",
|
||||
"payout_ratio", "fcf_dividend_cover", "free_cashflow",
|
||||
"total_cash_dividend",
|
||||
# 标量型质量指标。
|
||||
# 刻意**不含 ttm_dps** —— 它已由上面的序列循环按窗口产出(带完整分布统计),
|
||||
# 这里再产一次会写出两条 (ttm_dps, window=0) 行:落库时互相覆盖,
|
||||
# 而「哪条胜出」取决于写入顺序,属于不确定行为。
|
||||
#
|
||||
# 同理 **不含 free_cashflow**:``fin_latest`` 已按「最近一期公告」产出该代码,
|
||||
# 而这里是**与分红同一财年**的自由现金流 —— 两个不同口径不能共用一个代码。
|
||||
# 后者改用 ``dividend_fy_free_cashflow``,两者都保留、都唯一。
|
||||
for code, emit_code in (
|
||||
("dividend_continuity_years", None), ("dividend_years_in_window", None),
|
||||
("dps_cagr_5y", None), ("dps_volatility", None),
|
||||
("payout_ratio", None), ("fcf_dividend_cover", None),
|
||||
("free_cashflow", "dividend_fy_free_cashflow"),
|
||||
("total_cash_dividend", None),
|
||||
):
|
||||
v = _f(st.get(code))
|
||||
if v is None:
|
||||
continue
|
||||
out.append(self._stat_row(sym, code, None, {"n_obs": 1}, v, None, "OK"))
|
||||
out.append(self._stat_row(sym, emit_code or code, None, {"n_obs": 1}, v, None, "OK"))
|
||||
# DPS 年度序列(用于趋势展示与分位)
|
||||
dps_by_year = st.get("dps_by_year") or {}
|
||||
if dps_by_year:
|
||||
@@ -663,7 +729,16 @@ class ProfileBuilder:
|
||||
}
|
||||
|
||||
def _load_daily_basic(self, syms: list[str], start: date, end: date) -> pd.DataFrame:
|
||||
"""批量取每日指标(估值序列),按 symbol 分片以控制单条 SQL 的规模。"""
|
||||
"""批量取每日指标(估值序列),按 symbol 分片以控制单条 SQL 的规模。
|
||||
|
||||
**必须 ``ORDER BY symbol, trade_date``**:下游按「最后一个观测」取当日值
|
||||
(``current = sub.iloc[-1]``),若行序不是日期序就会取到任意一天的值。
|
||||
早期实现漏了排序 —— 在「行按日期顺序插入」时侥幸正确,
|
||||
但 2026-10-04 回补 2005-2014 时新行是**追加**进去的,
|
||||
于是同一股票的行序变成「2015-2026 在前、2005-2014 在后」,
|
||||
格力电器 2018-05-18 的 PE(TTM) 被取成 2014 年的 8.51(真值 12.12)。
|
||||
等价性测试(实时画像 vs 批量画像)当场抓到该分歧。
|
||||
"""
|
||||
out: list[pd.DataFrame] = []
|
||||
cfg = load_config("datasource")
|
||||
for i in range(0, len(syms), 500):
|
||||
@@ -673,7 +748,8 @@ class ProfileBuilder:
|
||||
params.update({f"s{j}": s for j, s in enumerate(batch)})
|
||||
df = db.read_sql(
|
||||
"SELECT symbol, trade_date, pe_ttm, pb, ps_ttm, dv_ttm "
|
||||
f"FROM daily_basic WHERE symbol IN ({ph}) AND trade_date BETWEEN :start AND :end",
|
||||
f"FROM daily_basic WHERE symbol IN ({ph}) AND trade_date BETWEEN :start AND :end "
|
||||
"ORDER BY symbol, trade_date",
|
||||
params, cfg=cfg,
|
||||
)
|
||||
if not df.empty:
|
||||
@@ -682,7 +758,8 @@ class ProfileBuilder:
|
||||
return pd.DataFrame(columns=["symbol", "trade_date", "pe_ttm", "pb", "ps_ttm"])
|
||||
df = pd.concat(out, ignore_index=True)
|
||||
df["trade_date"] = pd.to_datetime(df["trade_date"])
|
||||
return df
|
||||
# 分片拼接后仍需全局有序(各分片内部有序 ≠ 整体有序的日期序)
|
||||
return df.sort_values(["symbol", "trade_date"], ignore_index=True)
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 落库
|
||||
|
||||
@@ -0,0 +1,119 @@
|
||||
"""画像窗口的**实际覆盖度**(不是「报告了 N 年」,而是「真的有 N 年数据」)。
|
||||
|
||||
**为什么需要它**:``window_slice(asof, 5)`` 的语义是「把已有数据切成最近 5 年」,
|
||||
不是「保证有 5 年数据」。若数据起点晚于窗口左端,窗口会被**静默截短**,
|
||||
而 ``_stat_row`` 只看 ``n_obs >= min(min_obs_days, 20)`` 就标 ``OK`` ——
|
||||
20 个观测(约 1 个月)也算通过。
|
||||
|
||||
实测(600036.SH,``dv_yield``,5 年窗口):
|
||||
|
||||
| asof | 窗口内实际观测 | 应有权重 | 覆盖率 |
|
||||
|---|---:|---:|---:|
|
||||
| 2015-12-31 | 239 | 1212 | 19.7% |
|
||||
| 2016-12-30 | 483 | 1212 | 39.9% |
|
||||
| 2017-12-29 | 727 | 1212 | 60.0% |
|
||||
| 2018-12-28 | 970 | 1212 | 80.0% |
|
||||
| 2019-12-31 | 1214 | 1215 | 99.9% |
|
||||
| 2020 起 | ≈1215 | ≈1215 | 100% |
|
||||
|
||||
而 ``n_obs=817`` 的 2018-05-18、四个窗口(0/5/8/10)**报出完全相同的 n_obs**
|
||||
—— 这正是「被数据起点截断」的指纹。
|
||||
|
||||
本模块提供统一的分母(**交易日历的真实开市天数**,不是 243 这种近似),
|
||||
供实时画像(``profile.pit``)与批量画像(``profile.builder``)共用。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import date, timedelta
|
||||
from typing import Any
|
||||
|
||||
#: 覆盖率低于该值时,画像页与 CLI 打印警告(不改变任何判定,只是不让人误以为有 5 年)
|
||||
WARN_COVERAGE = 0.95
|
||||
|
||||
|
||||
def window_start(asof: date, years: int) -> date:
|
||||
"""窗口左端(与 ``factor.dividend_yield.window_slice`` 完全一致的口径)。"""
|
||||
return asof - timedelta(days=int(years * 365.25))
|
||||
|
||||
|
||||
def expected_trading_days(repo: Any, asof: date, years: int) -> int:
|
||||
"""``(asof - N 年, asof]`` 内**应有**的交易日数(按交易日历)。
|
||||
|
||||
``years <= 0`` 表示全历史:没有可比的「应有」天数,返回 0 让调用方跳过覆盖率判定。
|
||||
|
||||
区间必须是**左开右闭** —— 与 ``factor.dividend_yield.window_slice`` 的
|
||||
``trade_date > start`` 完全一致。``repo.trading_days`` 取的是闭区间,
|
||||
所以左端点恰为交易日时要减 1;否则覆盖率永远差一天、达不到 100%,
|
||||
会让 ``min_window_coverage = 1.0`` 变成「永远拒绝」。
|
||||
"""
|
||||
if years <= 0:
|
||||
return 0
|
||||
lo = window_start(asof, years)
|
||||
days = repo.trading_days(lo, asof)
|
||||
n = len(days)
|
||||
if n and days[0] == lo:
|
||||
n -= 1
|
||||
return n
|
||||
|
||||
|
||||
def coverage_ratio(n_obs: int, expected: int) -> float | None:
|
||||
"""实际观测数 / 应有交易日数,封顶 1.0。``expected<=0`` 时返回 None(不适用)。"""
|
||||
if expected <= 0:
|
||||
return None
|
||||
if n_obs >= expected:
|
||||
return 1.0
|
||||
return max(0.0, float(n_obs) / float(expected))
|
||||
|
||||
|
||||
def expected_by_window(repo: Any, asof: date, windows: list[int]) -> dict[int, int]:
|
||||
"""``{窗口年数: 应有交易日数}``(含 0 → 0)。"""
|
||||
return {int(w): expected_trading_days(repo, asof, int(w)) for w in windows}
|
||||
|
||||
|
||||
def summarise(stat_rows: list[dict[str, Any]], expected: dict[int, int]) -> dict[str, Any]:
|
||||
"""把一批 stat 行按窗口汇总覆盖率(供 CLI/报告打印警告)。
|
||||
|
||||
只统计 :data:`DAILY_OBSERVATION_METRICS` —— 其余指标的 ``n_obs`` 是财年数或 1,
|
||||
用交易日当分母会算出「0.4%」这种量纲错误的覆盖率。
|
||||
|
||||
返回 ``{"windows": {年数: {"min": 最低覆盖率, "n": 行数}}, "worst": (年数, 比率)}``。
|
||||
"""
|
||||
from hdiv.core.metrics import DAILY_OBSERVATION_METRICS
|
||||
|
||||
agg: dict[int, list[float]] = {}
|
||||
for r in stat_rows:
|
||||
if r.get("metric_code") not in DAILY_OBSERVATION_METRICS:
|
||||
continue
|
||||
wy = int(r.get("window_years") or 0)
|
||||
exp = expected.get(wy, 0)
|
||||
c = coverage_ratio(int(r.get("n_obs") or 0), exp)
|
||||
if c is None:
|
||||
continue
|
||||
agg.setdefault(wy, []).append(c)
|
||||
windows = {
|
||||
wy: {"min": min(vals), "n": len(vals)} for wy, vals in sorted(agg.items())
|
||||
}
|
||||
worst: tuple[int, float] | None = None
|
||||
for wy, info in windows.items():
|
||||
if worst is None or info["min"] < worst[1]:
|
||||
worst = (wy, info["min"])
|
||||
return {"windows": windows, "worst": worst}
|
||||
|
||||
|
||||
def format_warning(summary: dict[str, Any]) -> str | None:
|
||||
"""覆盖率不足时的一句话说明(否则 None)。
|
||||
|
||||
逐个列出**所有**不足的窗口,而不是只报最差的那个 ——
|
||||
否则「5 年已齐、只有 10 年不足」也会被说成「数据不足」,容易误导。
|
||||
"""
|
||||
windows = (summary or {}).get("windows") or {}
|
||||
short = [(wy, info["min"]) for wy, info in sorted(windows.items())
|
||||
if info["min"] < WARN_COVERAGE]
|
||||
if not short:
|
||||
return None
|
||||
detail = "、".join(f"{wy} 年窗口 {cov:.1%}" for wy, cov in short)
|
||||
return (
|
||||
f"窗口数据不足:{detail}"
|
||||
f"(窗口被数据起点截短,分位/统计量的实际样本期短于名义窗口)"
|
||||
)
|
||||
@@ -0,0 +1,555 @@
|
||||
"""Point-in-Time(实时)个股画像服务。
|
||||
|
||||
**为什么需要它**:``ProfileBuilder`` 是**批量 + 单一时点**的研究工具 ——
|
||||
它的 ``run()`` 一次算完整个股票池、落库一份快照,供人看。回测需要的却是
|
||||
「每个决策日按当时可见的数据重新画像」。两者共用同一套指标定义,但调用形态不同。
|
||||
|
||||
本模块提供回测侧的形态:
|
||||
|
||||
1. **PIT 语义与 ProfileBuilder 完全一致**:价格、每日指标、分红、财报一律只在
|
||||
``<= asof`` 的范围内取数;每个窗口用同一把 ``window_slice`` 切分。
|
||||
等价性有回归测试(``tests/test_profile_pit.py``),而不是靠注释保证。
|
||||
2. **惰性**:只在「买入条件已触发」时才计算 —— 绝大多数股票日根本不需要画像。
|
||||
3. **可复用**:跨决策日共享的面板(价格、每日指标、指数)只载入一次;
|
||||
与时点强相关的面板(分红、财报)按 asof 缓存,同一 asof 内多只股票复用。
|
||||
这正是「长期数据可以沿用、在触发条件时计算」的落地方式。
|
||||
4. **成本可观测**:``stats()`` 汇报载入次数/查询次数,避免「悄悄变慢」。
|
||||
|
||||
一次完整回测(12 年、每月评估、49 只候选)的画像计算量取决于触发次数,
|
||||
而不是 12 年 × 49 只 —— 这是惰性的核心收益。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
from datetime import date, timedelta
|
||||
from typing import Any
|
||||
|
||||
import pandas as pd
|
||||
|
||||
from hdiv.core.errors import HdivError
|
||||
from hdiv.core.metrics import (
|
||||
ALL_METRICS,
|
||||
DAILY_OBSERVATION_METRICS,
|
||||
DIVIDEND_METRICS,
|
||||
FINANCIAL_METRICS,
|
||||
GATE_METRICS,
|
||||
LIQUIDITY_METRICS,
|
||||
PERCENTILE_METRICS,
|
||||
RETURN_METRICS,
|
||||
VALUATION_METRICS,
|
||||
)
|
||||
from hdiv.data.repo import Repo
|
||||
from hdiv.factor.dividend_yield import build_dps_events
|
||||
from hdiv.profile.builder import METRIC_META, ProfileBuilder
|
||||
from hdiv.profile.coverage import coverage_ratio, expected_by_window
|
||||
|
||||
__all__ = [
|
||||
"ALL_METRICS", "DIVIDEND_METRICS", "FINANCIAL_METRICS", "GATE_METRICS",
|
||||
"LIQUIDITY_METRICS", "METRIC_META", "PERCENTILE_METRICS", "RETURN_METRICS",
|
||||
"VALUATION_METRICS", "PitProfileService", "ProfileSnapshot",
|
||||
"evaluate_gate", "metrics_needing_dividends", "metrics_needing_financials",
|
||||
]
|
||||
|
||||
# 指标分组定义在 hdiv.core.metrics(叶子模块),此处引用同一份 ——
|
||||
# 配置校验(core.config)与画像实现(profile.pit)不允许出现两个清单。
|
||||
|
||||
|
||||
def metrics_needing_financials(metrics: set[str]) -> bool:
|
||||
"""是否需要财报面板。
|
||||
|
||||
**分红质量指标也算「需要财报」**:``ProfileBuilder._dividend_quality_stats``
|
||||
用「最近一个已公告年报」的 ``fin_end_date`` 反推考核财年(``_target_years``),
|
||||
再用同一财年的 ``annual_financials`` 算支付率/FCF 覆盖。缺了 ``fin_latest``
|
||||
会静默退回 ``asof.year - 1`` 这个猜测值 —— 实测会让格力电器
|
||||
2018-05-18 的 ``dividend_continuity_years`` 从 10 变成 4。
|
||||
宁可多付一次查询,也不接受两套口径。
|
||||
"""
|
||||
return bool(metrics & (FINANCIAL_METRICS | DIVIDEND_METRICS))
|
||||
|
||||
|
||||
def metrics_needing_dividends(metrics: set[str]) -> bool:
|
||||
return bool(metrics & (DIVIDEND_METRICS | VALUATION_METRICS))
|
||||
|
||||
# 指标分组(VALUATION/RETURN/DIVIDEND/FINANCIAL/LIQUIDITY_METRICS、
|
||||
# PERCENTILE_METRICS、ALL_METRICS)定义在 ``hdiv.core.metrics`` 这个叶子模块里,
|
||||
# 此处已 import —— 配置校验(core.config)与画像实现共用同一份清单,
|
||||
# 不允许出现两个版本的「允许哪些指标」。
|
||||
|
||||
#: ``ProfileBuilder._profile_one`` 会直接对 ``fin_*`` 面板做 ``["symbol"]`` 取列,
|
||||
#: 所以「不需要财报」时也必须给出**带列名的空表**,而不是无列的 ``DataFrame()``
|
||||
#: (否则会 KeyError,而不是安静地跳过财务指标)。
|
||||
_EMPTY_FIN_HIST = (
|
||||
"symbol", "year", "roe", "roic", "grossprofit_margin", "netprofit_margin",
|
||||
"ocf_to_profit", "ocf_to_netprofit_calc",
|
||||
)
|
||||
_EMPTY_FIN_AVG = (
|
||||
"symbol", "roe_avg", "roic_avg", "gross_margin_avg", "net_margin_avg",
|
||||
"ocf_to_profit_avg", "fin_years_count", "fin_latest_year",
|
||||
)
|
||||
_EMPTY_FIN_LATEST = (
|
||||
"symbol", "end_date", "ann_date", "debt_ratio", "free_cashflow",
|
||||
"n_income_attr_p", "total_assets",
|
||||
)
|
||||
_EMPTY_BASICS = ("symbol", "trade_date", "pe_ttm", "pb", "ps_ttm")
|
||||
|
||||
|
||||
def _with_columns(df: pd.DataFrame, cols: tuple[str, ...]) -> pd.DataFrame:
|
||||
"""空表 → 带列名的空表;非空表原样返回。"""
|
||||
if df.empty and "symbol" not in df.columns:
|
||||
return pd.DataFrame(columns=list(cols))
|
||||
return df
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 单只股票的画像快照
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@dataclass
|
||||
class ProfileSnapshot:
|
||||
"""某只股票在某个 asof 的实时画像(只保留决策需要的形态)。"""
|
||||
|
||||
symbol: str
|
||||
asof: date
|
||||
window_years: int
|
||||
#: 指标 → 当前值(``current_value``)
|
||||
values: dict[str, float] = field(default_factory=dict)
|
||||
#: 指标 → 当前在窗口分布中的分位(仅 PERCENTILE_METRICS)
|
||||
percentiles: dict[str, float] = field(default_factory=dict)
|
||||
#: 指标 → ``OK`` / ``INSUFFICIENT``(样本不足时不得当作可用)
|
||||
status: dict[str, str] = field(default_factory=dict)
|
||||
#: 指标 → 实际取数的窗口年数(0 = 全历史)
|
||||
windows: dict[str, int] = field(default_factory=dict)
|
||||
#: 指标 → 窗口内的实际观测数
|
||||
n_obs: dict[str, int] = field(default_factory=dict)
|
||||
#: 指标 → 窗口**实际覆盖率**(1.0 = 名义窗口被完整覆盖;窗口 0 不适用)
|
||||
coverage: dict[str, float] = field(default_factory=dict)
|
||||
#: 安全边际分项得分(供留痕,不参与闸门判定)
|
||||
scores: dict[str, float] = field(default_factory=dict)
|
||||
|
||||
def get(self, metric: str) -> tuple[float | None, str, int]:
|
||||
"""返回 ``(值, 状态, 窗口)``。缺失指标的状态为 ``MISSING``。"""
|
||||
return (
|
||||
self.values.get(metric),
|
||||
self.status.get(metric, "MISSING"),
|
||||
self.windows.get(metric, -1),
|
||||
)
|
||||
|
||||
def get_percentile(self, metric: str) -> float | None:
|
||||
return self.percentiles.get(metric)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 按 asof 缓存的时点上下文
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@dataclass
|
||||
class _AsOfContext:
|
||||
"""同一 asof 上所有股票共用的 PIT 面板(惰性装载)。"""
|
||||
|
||||
asof: date
|
||||
symbols: list[str]
|
||||
dividends: pd.DataFrame
|
||||
div_by_symbol: dict[str, list[dict[str, Any]]]
|
||||
fin_hist: pd.DataFrame = field(default_factory=pd.DataFrame)
|
||||
fin_avg: pd.DataFrame = field(default_factory=pd.DataFrame)
|
||||
fin_latest: pd.DataFrame = field(default_factory=pd.DataFrame)
|
||||
fy_table: dict[tuple[str, int], Any] = field(default_factory=dict)
|
||||
needs_financial: bool = False
|
||||
#: ``{窗口年数: 该窗口应有的交易日数}`` —— 覆盖率的分母(按交易日历,非近似)
|
||||
expected_obs: dict[int, int] = field(default_factory=dict)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 服务
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
class PitProfileService:
|
||||
"""实时画像服务(PIT,惰性,按 asof 缓存)。
|
||||
|
||||
用法::
|
||||
|
||||
svc = PitProfileService(window_years=5)
|
||||
svc.prepare(symbols, data_start, end)
|
||||
snap = svc.snapshot("600036.SH", date(2018, 5, 18))
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
window_years: int = 5,
|
||||
repo: Repo | None = None,
|
||||
builder: ProfileBuilder | None = None,
|
||||
load_index: bool = True,
|
||||
load_basics: bool = True,
|
||||
) -> None:
|
||||
self.window_years = int(window_years)
|
||||
self.repo = repo or Repo()
|
||||
self.builder = builder or ProfileBuilder.from_config()
|
||||
self.max_years = max(self.builder.config.windows_years)
|
||||
if self.window_years != 0 and self.window_years not in self.builder.config.windows_years:
|
||||
raise HdivError(
|
||||
f"实时画像窗口 {self.window_years} 年不可用:config/profile.yml 的 "
|
||||
f"windows_years={self.builder.config.windows_years} 里没有它。\n"
|
||||
f" 请把它加进 windows_years(例如 [3, 5, 8, 10]),"
|
||||
f"或把策略的 profile_gate.window_years 改成其中一个值。"
|
||||
)
|
||||
self.load_index = load_index
|
||||
self.load_basics = load_basics
|
||||
self._prepared = False
|
||||
self._symbols: list[str] = []
|
||||
self._price = pd.DataFrame()
|
||||
self._basics = pd.DataFrame()
|
||||
self._index = pd.DataFrame()
|
||||
self._all_dividends = pd.DataFrame()
|
||||
self._ctx: dict[date, _AsOfContext] = {}
|
||||
self._snapshots: dict[tuple[str, date], ProfileSnapshot] = {}
|
||||
#: 闸门规则用到的指标集合;None = 未知,按「全都可能需要」处理(保守)
|
||||
self._needed: set[str] | None = None
|
||||
self._financial_required: bool | None = None
|
||||
self._counters: dict[str, int] = {
|
||||
"price_loaded": 0, "basics_loaded": 0, "asof_contexts": 0,
|
||||
"financial_loads": 0, "liquidity_loads": 0,
|
||||
"snapshots_computed": 0, "snapshots_cached": 0,
|
||||
}
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 成本控制:只载入规则真正需要的面板
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
def configure(self, metrics: set[str]) -> None:
|
||||
"""声明闸门用到的指标集合。
|
||||
|
||||
这一步是**成本控制的关键**:若规则里没有任何财务指标,就不必付财报全表
|
||||
查询的代价(各约 30 万行)。未调用时按「全都可能需要」保守处理。
|
||||
"""
|
||||
self._needed = set(metrics)
|
||||
self._financial_required = metrics_needing_financials(self._needed)
|
||||
|
||||
def _needs(self, metric: str) -> bool:
|
||||
return self._needed is None or metric in self._needed
|
||||
|
||||
def _needs_financials(self) -> bool:
|
||||
"""未 ``configure`` 时按「需要」处理 —— 宁可多查,不可少算。"""
|
||||
return self._financial_required is not False
|
||||
|
||||
def require_financials(self, required: bool) -> None:
|
||||
"""显式声明是否需要财报面板(覆盖 ``configure`` 的推断)。"""
|
||||
self._financial_required = bool(required)
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 批量预载(跨越整个回测区间、与 asof 无关的部分)
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
def prepare(self, symbols: list[str], start: date, end: date) -> None:
|
||||
"""载入跨决策日共享的面板。
|
||||
|
||||
只有 ``trade_date`` 范围过滤,没有 PIT 语义 —— 真正的 PIT 剪裁发生在
|
||||
:meth:`snapshot` 里逐 asof 进行(与 ``ProfileBuilder.run`` 的取数起点
|
||||
规则完全一致:``asof.year - max_years - 1`` 的 1 月 1 日)。
|
||||
"""
|
||||
self._symbols = sorted(set(symbols))
|
||||
if not self._symbols:
|
||||
self._prepared = True
|
||||
return
|
||||
self._price = self.repo.price_history(self._symbols, start, end, adjust="none")
|
||||
self._counters["price_loaded"] += 1
|
||||
if self.load_basics:
|
||||
self._basics = self.builder._load_daily_basic(self._symbols, start, end)
|
||||
self._counters["basics_loaded"] += 1
|
||||
if self.load_index:
|
||||
self._index = self.repo.index_history("000300.SH", start, end)
|
||||
# 分红:为**整个回测区间**取一次超集,逐 asof 再用与 repo.dividend_records
|
||||
# 完全相同的三重 PIT 条件(imp_ann_date / ex_date / 回看窗口)在 pandas 里剪裁。
|
||||
#
|
||||
# 超集下界必须覆盖**最早**的 asof 所需的回看窗口,而不是最后一个 asof ——
|
||||
# 早期实现按 end 回看 13 年,于是 2018 年的画像拿不到 2006-2012 的分红,
|
||||
# 格力电器的 dividend_continuity_years 被算成 4(真值 10)。
|
||||
span_start = start - timedelta(days=int((self.max_years + 2) * 365.25))
|
||||
span_years = int((end - span_start).days / 365.25) + 2
|
||||
self._all_dividends = self.repo.dividend_records(end, years_back=span_years)
|
||||
if not self._all_dividends.empty:
|
||||
self._all_dividends = self._all_dividends[
|
||||
self._all_dividends["symbol"].isin(set(self._symbols))
|
||||
]
|
||||
self._prepared = True
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 单股快照
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
def snapshot(self, symbol: str, asof: date) -> ProfileSnapshot | None:
|
||||
"""计算(或取缓存)``symbol`` 在 ``asof`` 的实时画像。"""
|
||||
if not self._prepared:
|
||||
raise HdivError("PitProfileService 必须先 prepare(symbols, start, end)")
|
||||
key = (symbol, asof)
|
||||
hit = self._snapshots.get(key)
|
||||
if hit is not None:
|
||||
self._counters["snapshots_cached"] += 1
|
||||
return hit
|
||||
ctx = self._context(asof, needs_financial=self._needs_financials())
|
||||
snap = self._compute(symbol, asof, ctx)
|
||||
if snap is not None:
|
||||
self._snapshots[key] = snap
|
||||
self._counters["snapshots_computed"] += 1
|
||||
return snap
|
||||
|
||||
def stats(self) -> dict[str, int]:
|
||||
out = dict(self._counters)
|
||||
out["distinct_asof"] = len(self._ctx)
|
||||
out["cached_symbols"] = len(self._snapshots)
|
||||
return out
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 内部
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
def _context(self, asof: date, *, needs_financial: bool | None) -> _AsOfContext:
|
||||
hit = self._ctx.get(asof)
|
||||
if hit is not None and (hit.needs_financial or not needs_financial):
|
||||
return hit
|
||||
# ---- 分红:与 repo.dividend_records 完全相同的 PIT 三重条件 ----
|
||||
d = self._all_dividends
|
||||
if d.empty:
|
||||
div = d
|
||||
else:
|
||||
since = asof - timedelta(days=int((self.max_years + 2) * 365.25))
|
||||
imp = pd.to_datetime(d["imp_ann_date"]).dt.date
|
||||
ex = pd.to_datetime(d["ex_date"]).dt.date
|
||||
div = d[(imp <= asof) & (ex <= asof) & (ex >= since)]
|
||||
div_by_symbol: dict[str, list[dict[str, Any]]] = {}
|
||||
if not div.empty:
|
||||
for rec in div.to_dict("records"):
|
||||
div_by_symbol.setdefault(rec["symbol"], []).append(rec)
|
||||
|
||||
ctx = _AsOfContext(
|
||||
asof=asof, symbols=self._symbols, dividends=div,
|
||||
div_by_symbol=div_by_symbol, needs_financial=bool(needs_financial),
|
||||
expected_obs=expected_by_window(
|
||||
self.repo, asof,
|
||||
# 窗口不止 profile.yml 的 [5,8,10]:收益类指标会产出 1/3 年窗口
|
||||
# (ret_1y / ret_3y / max_drawdown_3y),分母必须一并备好
|
||||
sorted({int(w) for w in self.builder.config.windows_years} | {1, 3}),
|
||||
),
|
||||
)
|
||||
if needs_financial:
|
||||
syms = self._symbols
|
||||
ctx.fin_hist = self.repo.annual_financial_history(
|
||||
asof, years=self.max_years + 1, symbols=syms
|
||||
)
|
||||
ctx.fin_avg = self.repo.annual_financial_averages(
|
||||
asof, years=5, symbols=syms, hist=ctx.fin_hist
|
||||
)
|
||||
ctx.fin_latest = self.repo.financial_panel(asof, symbols=syms)
|
||||
annual = self.repo.annual_financials(asof, years=self.max_years + 2, symbols=syms)
|
||||
if not annual.empty:
|
||||
ctx.fy_table = {
|
||||
(r["symbol"], int(r["year"])): r for _, r in annual.iterrows()
|
||||
}
|
||||
self._counters["financial_loads"] += 1
|
||||
self._ctx[asof] = ctx
|
||||
self._counters["asof_contexts"] += 1
|
||||
return ctx
|
||||
|
||||
def _compute(
|
||||
self, symbol: str, asof: date, ctx: _AsOfContext
|
||||
) -> ProfileSnapshot | None:
|
||||
"""调用 ``ProfileBuilder._profile_one`` —— **指标定义的单一口径来源**。"""
|
||||
start = date(asof.year - self.max_years - 1, 1, 1)
|
||||
|
||||
def _slice(df: pd.DataFrame) -> pd.DataFrame:
|
||||
if df.empty:
|
||||
return df
|
||||
td = pd.to_datetime(df["trade_date"]).dt.date
|
||||
return df[(td >= start) & (td <= asof)]
|
||||
|
||||
price = _slice(self._price) if not self._price.empty else self._price
|
||||
price = price[price["symbol"] == symbol] if not price.empty else price
|
||||
if price.empty:
|
||||
return None
|
||||
basics = _slice(self._basics) if not self._basics.empty else self._basics
|
||||
if not basics.empty:
|
||||
basics = basics[basics["symbol"] == symbol]
|
||||
basics = _with_columns(basics, _EMPTY_BASICS)
|
||||
index = _slice(self._index) if not self._index.empty else self._index
|
||||
sym_div = ctx.div_by_symbol.get(symbol, [])
|
||||
# ProfileBuilder 期望的 events 是 {symbol: DataFrame}
|
||||
events = build_dps_events(pd.DataFrame(sym_div)) if sym_div else {}
|
||||
# 流动性是逐 (股票, 时点) 的查询:只有在规则真的用到时才付出这个代价
|
||||
if self._needs("avg_amount_20d"):
|
||||
avg = self.repo.avg_amount(asof, window=20, symbols=[symbol])
|
||||
self._counters["liquidity_loads"] += 1
|
||||
else:
|
||||
avg = pd.DataFrame(columns=["symbol", "avg_amount", "n"])
|
||||
|
||||
res = self.builder._profile_one(
|
||||
symbol, asof, price, events, basics, index, avg,
|
||||
_with_columns(ctx.fin_hist, _EMPTY_FIN_HIST),
|
||||
_with_columns(ctx.fin_avg, _EMPTY_FIN_AVG),
|
||||
_with_columns(ctx.fin_latest, _EMPTY_FIN_LATEST),
|
||||
ctx.div_by_symbol, ctx.fy_table,
|
||||
# 闸门只需要当日值与分位,不需要绘图序列(省掉约一半耗时)
|
||||
build_series=False,
|
||||
)
|
||||
if res is None:
|
||||
return None
|
||||
|
||||
snap = ProfileSnapshot(symbol=symbol, asof=asof, window_years=self.window_years)
|
||||
by_code: dict[str, dict[str, Any]] = {}
|
||||
for row in res["stats"]:
|
||||
code = row["metric_code"]
|
||||
wy = int(row["window_years"])
|
||||
# 序列型指标同时有「全历史(0)」与各窗口行 —— 按配置的窗口优先,
|
||||
# 找不到则退回全历史,并把实际窗口如实记录(windows[code])。
|
||||
if code not in by_code or self._prefer(wy, by_code[code]["window_years"]):
|
||||
by_code[code] = row
|
||||
for code, row in by_code.items():
|
||||
cur = row.get("current_value")
|
||||
if cur is not None:
|
||||
snap.values[code] = float(cur)
|
||||
pct = row.get("current_percentile")
|
||||
if pct is not None:
|
||||
snap.percentiles[code] = float(pct)
|
||||
snap.status[code] = str(row.get("status") or "MISSING")
|
||||
snap.windows[code] = int(row["window_years"])
|
||||
n_obs = int(row.get("n_obs") or 0)
|
||||
snap.n_obs[code] = n_obs
|
||||
# 覆盖率:实际观测数 / 该窗口应有的交易日数。
|
||||
# **这是「名义 5 年」与「真的有 5 年数据」的区别所在**。
|
||||
# 只对**观测单位是交易日**的指标计算 —— 年报均值类的 n_obs 是财年数、
|
||||
# 标量类的 n_obs 是 1,拿交易日当分母是量纲错误。
|
||||
# 窗口 0(全历史)没有「应有」天数,记为 1.0(不参与判定)。
|
||||
if code in DAILY_OBSERVATION_METRICS:
|
||||
cov = coverage_ratio(
|
||||
n_obs, ctx.expected_obs.get(int(row["window_years"]), 0)
|
||||
)
|
||||
else:
|
||||
cov = None
|
||||
snap.coverage[code] = 1.0 if cov is None else cov
|
||||
for row in res["scores"]:
|
||||
if row.get("score") is not None:
|
||||
snap.scores[row["score_code"]] = float(row["score"])
|
||||
return snap
|
||||
|
||||
def _prefer(self, new_wy: int, old_wy: int) -> bool:
|
||||
"""窗口优先级:配置窗口 > 全历史 > 其它;**同窗口时后来者胜出**。
|
||||
|
||||
「同窗口后来者胜出」是刻意与落库语义对齐:``hd_profile_stat`` 对
|
||||
``(run_id, symbol, metric_code, window_years)`` 唯一,若画像产出了重复行,
|
||||
数据库里留下的是**最后写入**的那条。实时画像必须与页面上看到的数是同一个。
|
||||
(重复行本身正在被逐一消除,这里只是兜底,不允许出现两套口径。)
|
||||
"""
|
||||
rank = {self.window_years: 0, 0: 1}
|
||||
rn, ro = rank.get(new_wy, 2), rank.get(old_wy, 2)
|
||||
return rn <= ro
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 闸门判定
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def evaluate_gate(
|
||||
rules: list[dict[str, Any]],
|
||||
snapshot: ProfileSnapshot | None,
|
||||
*,
|
||||
on_unverifiable: str = "reject",
|
||||
min_window_coverage: float = 0.0,
|
||||
) -> dict[str, Any]:
|
||||
"""按规则逐条判定;返回可直接写入 ``reason_json`` 的可追溯结果。
|
||||
|
||||
三种结局:
|
||||
|
||||
- ``PASS`` —— 全部规则成立
|
||||
- ``REJECT`` —— 至少一条规则不成立
|
||||
- ``UNVERIFIABLE`` —— 指标缺失、样本不足,**或窗口数据覆盖不足**;
|
||||
按 ``on_unverifiable`` 决定是保守淘汰(``reject``)还是放行(``pass``)
|
||||
|
||||
**不猜**:指标缺失(MISSING)、样本不足(INSUFFICIENT)绝不当作 0 或当作通过。
|
||||
|
||||
``min_window_coverage``:窗口实际覆盖率下限(1.0 = 必须完整覆盖名义窗口)。
|
||||
默认 0 表示**不因覆盖率淘汰**(保持改造前行为);设成 1.0 时,
|
||||
「名义 5 年但实际只有 3.4 年数据」会被判为无法验证。
|
||||
"""
|
||||
checks: list[dict[str, Any]] = []
|
||||
failed: list[str] = []
|
||||
unverifiable: list[str] = []
|
||||
|
||||
for r in rules:
|
||||
metric = str(r["metric"])
|
||||
stat = str(r.get("stat", "current_value"))
|
||||
op = str(r.get("op", ">="))
|
||||
threshold = float(r["value"])
|
||||
actual: float | None = None
|
||||
status = "MISSING"
|
||||
window = -1
|
||||
coverage = 1.0
|
||||
n_obs = 0
|
||||
if snapshot is not None:
|
||||
if stat == "current_percentile":
|
||||
actual = snapshot.get_percentile(metric)
|
||||
status = snapshot.status.get(metric, "MISSING")
|
||||
window = snapshot.windows.get(metric, -1)
|
||||
else:
|
||||
actual, status, window = snapshot.get(metric)
|
||||
coverage = snapshot.coverage.get(metric, 1.0)
|
||||
n_obs = snapshot.n_obs.get(metric, 0)
|
||||
# 覆盖率不足 = 用**不完整**的窗口算出来的统计量,不能当作已验证
|
||||
short_window = window > 0 and coverage < min_window_coverage - 1e-9
|
||||
ok: bool | None
|
||||
if actual is None or status != "OK":
|
||||
ok = None
|
||||
unverifiable.append(f"{metric}.{stat}")
|
||||
elif short_window:
|
||||
ok = None
|
||||
unverifiable.append(f"{metric}.window_coverage={coverage:.0%}")
|
||||
else:
|
||||
ok = _compare(actual, op, threshold)
|
||||
if not ok:
|
||||
failed.append(f"{metric}.{stat}{op}{threshold:g}")
|
||||
checks.append({
|
||||
"metric": metric, "stat": stat, "op": op, "threshold": threshold,
|
||||
"actual": actual, "status": status, "window_years": window,
|
||||
"n_obs": n_obs, "window_coverage": round(coverage, 4),
|
||||
"passed": ok,
|
||||
})
|
||||
|
||||
if failed:
|
||||
verdict = "REJECT"
|
||||
elif unverifiable:
|
||||
verdict = "REJECT" if on_unverifiable == "reject" else "PASS"
|
||||
else:
|
||||
verdict = "PASS"
|
||||
|
||||
detail: dict[str, Any] = {
|
||||
"verdict": verdict,
|
||||
"checks": checks,
|
||||
"failed": failed,
|
||||
"unverifiable": unverifiable,
|
||||
"on_unverifiable": on_unverifiable,
|
||||
"min_window_coverage": min_window_coverage,
|
||||
}
|
||||
if snapshot is not None:
|
||||
detail["asof"] = str(snapshot.asof)
|
||||
detail["window_years"] = snapshot.window_years
|
||||
if snapshot.scores:
|
||||
detail["safety_margin_scores"] = snapshot.scores
|
||||
return detail
|
||||
|
||||
|
||||
_OPS = {
|
||||
">=": lambda a, b: a >= b,
|
||||
"<=": lambda a, b: a <= b,
|
||||
">": lambda a, b: a > b,
|
||||
"<": lambda a, b: a < b,
|
||||
}
|
||||
|
||||
|
||||
def _compare(actual: float, op: str, threshold: float) -> bool:
|
||||
fn = _OPS.get(op)
|
||||
if fn is None: # pragma: no cover - 配置层已校验
|
||||
raise HdivError(f"不支持的比较符:{op}")
|
||||
return bool(fn(actual, threshold))
|
||||
+251
-5
@@ -21,6 +21,18 @@ import numpy as np
|
||||
import pandas as pd
|
||||
|
||||
from hdiv.core.config import load_config
|
||||
from hdiv.core.errors import HdivError
|
||||
|
||||
from hdiv.report.format import NumFmt
|
||||
|
||||
#: 净值曲线默认叠加的指数(沪深300)
|
||||
DEFAULT_INDEX_CODE = "000300.SH"
|
||||
|
||||
|
||||
def _fmt() -> NumFmt:
|
||||
"""当前配置的格式化器(每次读取,保证改配置立即生效)。"""
|
||||
return NumFmt.from_config()
|
||||
|
||||
from hdiv.data import db
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -151,7 +163,7 @@ def _yi(v: Any) -> str:
|
||||
|
||||
def _pct(v: Any) -> str:
|
||||
n = _num(v)
|
||||
return "—" if n is None else f"{n * 100:.2f}%"
|
||||
return "—" if n is None else _fmt().pct(n)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -550,6 +562,175 @@ def get_backtest(run_id: str) -> dict[str, Any] | None:
|
||||
return r
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Walk-forward(样本外验证)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def list_walkforwards() -> list[dict[str, Any]]:
|
||||
"""Walk-forward 运行列表。每个 wf_id 是一条独立记录。"""
|
||||
cfg = load_config("datasource")
|
||||
df = db.read_sql(
|
||||
"""
|
||||
SELECT w.wf_id, w.strategy_id, w.strategy_version, w.scheme,
|
||||
w.train_years, w.test_years, w.step_months, w.window_count,
|
||||
w.config_hash, w.data_version, w.code_version, w.status,
|
||||
w.created_at, w.config_json,
|
||||
(SELECT COUNT(*) FROM hd_walkforward_window k
|
||||
WHERE k.wf_id = w.wf_id) AS windows_actual
|
||||
FROM hd_walkforward_run w
|
||||
ORDER BY w.created_at DESC
|
||||
""",
|
||||
cfg=cfg,
|
||||
)
|
||||
out: list[dict[str, Any]] = []
|
||||
for _, row in df.iterrows():
|
||||
r = _rec(row)
|
||||
r["strategy"] = describe_strategy(_json_field(r.pop("config_json", None)) or {})
|
||||
r["title"] = (
|
||||
f"{r['strategy_id']} v{r['strategy_version']} · "
|
||||
f"{r['scheme']} 训练{r['train_years']}年/测试{r['test_years']}年 · "
|
||||
f"{r['window_count']} 窗口"
|
||||
)
|
||||
# 汇总样本外表现。逐窗口取「该窗口 test_run 的 total_return」,
|
||||
# 每个窗口只贡献一个样本 —— 早期版本靠 metric 行数反推窗口数,
|
||||
# 一旦某个指标缺失就会算错胜率。
|
||||
agg = db.read_sql(
|
||||
"""
|
||||
SELECT k.window_index,
|
||||
MAX(CASE WHEN m.metric_code='total_return'
|
||||
THEN m.metric_value END) AS ret,
|
||||
MAX(CASE WHEN m.metric_code='max_drawdown'
|
||||
THEN m.metric_value END) AS dd
|
||||
FROM hd_walkforward_window k
|
||||
LEFT JOIN hd_backtest_metric m
|
||||
ON m.run_id = k.test_run_id AND m.scope = 'all'
|
||||
AND (m.benchmark_code IS NULL OR m.benchmark_code = '')
|
||||
WHERE k.wf_id = :w
|
||||
GROUP BY k.window_index
|
||||
""",
|
||||
{"w": r["wf_id"]}, cfg=cfg,
|
||||
)
|
||||
rets = [x for x in agg["ret"].tolist() if x is not None]
|
||||
dds = [x for x in agg["dd"].tolist() if x is not None]
|
||||
r["oos"] = {
|
||||
"window_count": len(agg),
|
||||
"sample_count": len(rets),
|
||||
"mean_return": (sum(rets) / len(rets)) if rets else None,
|
||||
"win_rate": (sum(1 for x in rets if x > 0) / len(rets)) if rets else None,
|
||||
"worst_drawdown": min(dds) if dds else None,
|
||||
}
|
||||
out.append(r)
|
||||
return out
|
||||
|
||||
|
||||
def get_walkforward(wf_id: str) -> dict[str, Any] | None:
|
||||
"""Walk-forward 详情:逐窗口的样本内/外表现与冻结参数。"""
|
||||
cfg = load_config("datasource")
|
||||
head = db.read_sql(
|
||||
"SELECT * FROM hd_walkforward_run WHERE wf_id = :w", {"w": wf_id}, cfg=cfg
|
||||
)
|
||||
if head.empty:
|
||||
return None
|
||||
r = _rec(head.iloc[0])
|
||||
r["strategy"] = describe_strategy(_json_field(r.pop("config_json", None)) or {})
|
||||
r["config"] = _json_field(r.pop("config_json", None))
|
||||
r["title"] = (
|
||||
f"{r['strategy_id']} v{r['strategy_version']} · "
|
||||
f"{r['window_count']} 窗口({r['scheme']})"
|
||||
)
|
||||
|
||||
win = db.read_sql(
|
||||
"SELECT window_index, train_start, train_end, test_start, test_end, "
|
||||
" frozen_params_json, train_run_id, test_run_id "
|
||||
"FROM hd_walkforward_window WHERE wf_id = :w ORDER BY window_index",
|
||||
{"w": wf_id}, cfg=cfg,
|
||||
)
|
||||
run_ids = [x for x in win["train_run_id"].tolist() + win["test_run_id"].tolist() if x]
|
||||
mets: dict[str, dict[str, Any]] = {}
|
||||
if run_ids:
|
||||
placeholders = ",".join(f":r{i}" for i in range(len(run_ids)))
|
||||
params = {f"r{i}": v for i, v in enumerate(run_ids)}
|
||||
md = db.read_sql(
|
||||
f"SELECT run_id, metric_code, metric_value FROM hd_backtest_metric "
|
||||
f"WHERE run_id IN ({placeholders}) AND scope='all' "
|
||||
f" AND (benchmark_code IS NULL OR benchmark_code = '') "
|
||||
f" AND metric_code IN ('total_return','cagr','max_drawdown','sharpe',"
|
||||
f" 'trade_count','annual_volatility')",
|
||||
params, cfg=cfg,
|
||||
)
|
||||
for _, x in md.iterrows():
|
||||
mets.setdefault(x["run_id"], {})[x["metric_code"]] = _num(x["metric_value"])
|
||||
|
||||
# 基准指标存成 benchmark_<code>(如 benchmark_000300.SH),
|
||||
# 上面的查询用 benchmark_code='' 过滤掉了它 —— 于是页面上看不到
|
||||
# **超额收益**,而那恰恰是判读样本外最该看的数字。这里单独取回。
|
||||
bm = db.read_sql(
|
||||
f"SELECT run_id, metric_code, metric_value FROM hd_backtest_metric "
|
||||
f"WHERE run_id IN ({placeholders}) AND category = 'benchmark'",
|
||||
params, cfg=cfg,
|
||||
)
|
||||
for _, x in bm.iterrows():
|
||||
code = str(x["metric_code"]).replace("benchmark_", "")
|
||||
mets.setdefault(x["run_id"], {})[f"benchmark::{code}"] = _num(x["metric_value"])
|
||||
|
||||
# 基准:从 equity 表取被比较基准区间收益(回测落库时已算入 metric)
|
||||
windows = []
|
||||
for _, w in win.iterrows():
|
||||
tr, te = w["train_run_id"], w["test_run_id"]
|
||||
oos = {k: v for k, v in mets.get(te, {}).items() if not k.startswith("benchmark::")}
|
||||
bench = {k.split("::", 1)[1]: v
|
||||
for k, v in mets.get(te, {}).items() if k.startswith("benchmark::")}
|
||||
# 超额 = 策略样本外收益 − 基准同期收益(取第一个基准,与 backtest.yml 顺序一致)
|
||||
bench_code = next(iter(bench), None)
|
||||
bench_ret = bench.get(bench_code) if bench_code else None
|
||||
strat_ret = oos.get("total_return")
|
||||
windows.append({
|
||||
"window_index": int(w["window_index"]),
|
||||
"train_start": _v(w["train_start"]), "train_end": _v(w["train_end"]),
|
||||
"test_start": _v(w["test_start"]), "test_end": _v(w["test_end"]),
|
||||
"frozen_params": _json_field(w["frozen_params_json"]) or {},
|
||||
"train_run_id": tr, "test_run_id": te,
|
||||
"in_sample": {k: v for k, v in mets.get(tr, {}).items()
|
||||
if not k.startswith("benchmark::")},
|
||||
"out_of_sample": oos,
|
||||
"benchmark_code": bench_code,
|
||||
"benchmark_return": bench_ret,
|
||||
"excess_return": (strat_ret - bench_ret)
|
||||
if (strat_ret is not None and bench_ret is not None) else None,
|
||||
})
|
||||
|
||||
oos = [x["out_of_sample"].get("total_return") for x in windows
|
||||
if x["out_of_sample"].get("total_return") is not None]
|
||||
dds = [x["out_of_sample"].get("max_drawdown") for x in windows
|
||||
if x["out_of_sample"].get("max_drawdown") is not None]
|
||||
bench = [x["benchmark_return"] for x in windows if x["benchmark_return"] is not None]
|
||||
excess = [x["excess_return"] for x in windows if x["excess_return"] is not None]
|
||||
import statistics as _st
|
||||
|
||||
summary = {
|
||||
"window_count": len(windows),
|
||||
"oos_returns": oos,
|
||||
"benchmark_returns": bench,
|
||||
"excess_returns": excess,
|
||||
"benchmark_mean": (_st.fmean(bench) if bench else None),
|
||||
"excess_mean": (_st.fmean(excess) if excess else None),
|
||||
"excess_win_rate": (sum(1 for x in excess if x > 0) / len(excess)) if excess else None,
|
||||
"oos_mean": (_st.fmean(oos) if oos else None),
|
||||
"oos_median": (_st.median(oos) if oos else None),
|
||||
"oos_win_rate": (sum(1 for x in oos if x > 0) / len(oos)) if oos else None,
|
||||
"oos_worst_drawdown": (min(dds) if dds else None),
|
||||
# 稳定性 = 均值 / 标准差:<1 说明窗口间差异大于均值本身,结论不稳
|
||||
"oos_stability": (
|
||||
_st.fmean(oos) / _st.pstdev(oos)
|
||||
if len(oos) > 1 and _st.pstdev(oos) > 0 else None
|
||||
),
|
||||
}
|
||||
r["windows"] = windows
|
||||
r["summary"] = summary
|
||||
return r
|
||||
|
||||
|
||||
def get_backtest_metrics(run_id: str) -> list[dict[str, Any]]:
|
||||
cfg = load_config("datasource")
|
||||
return _records(db.read_sql(
|
||||
@@ -559,7 +740,30 @@ def get_backtest_metrics(run_id: str) -> list[dict[str, Any]]:
|
||||
))
|
||||
|
||||
|
||||
def get_backtest_equity(run_id: str) -> dict[str, Any]:
|
||||
def list_indices() -> list[dict[str, Any]]:
|
||||
"""可叠加到净值曲线上的基准指数(库里有多少列多少)。
|
||||
|
||||
一并返回各自的行情的起止日期与点数:前端据此提示
|
||||
「该指数在本次回测区间内没有行情」,而不是画一条空线让人猜。
|
||||
"""
|
||||
cfg = load_config("datasource")
|
||||
df = db.read_sql(
|
||||
"SELECT index_code, MAX(index_name) AS index_name, COUNT(*) AS points, "
|
||||
" MIN(trade_date) AS start, MAX(trade_date) AS end "
|
||||
"FROM hd_index_daily GROUP BY index_code ORDER BY index_code",
|
||||
cfg=cfg,
|
||||
)
|
||||
return [{
|
||||
"code": r["index_code"],
|
||||
"name": _v(r["index_name"]) or r["index_code"],
|
||||
"points": int(r["points"]),
|
||||
"start": _v(r["start"]),
|
||||
"end": _v(r["end"]),
|
||||
"is_default": r["index_code"] == DEFAULT_INDEX_CODE,
|
||||
} for _, r in df.iterrows()]
|
||||
|
||||
|
||||
def get_backtest_equity(run_id: str, *, index_code: str | None = None) -> dict[str, Any]:
|
||||
cfg = load_config("datasource")
|
||||
df = db.read_sql(
|
||||
"SELECT trade_date, nav, total_value, cash, position_value, drawdown, "
|
||||
@@ -568,13 +772,52 @@ def get_backtest_equity(run_id: str) -> dict[str, Any]:
|
||||
{"r": run_id}, cfg=cfg,
|
||||
)
|
||||
if df.empty:
|
||||
return {"dates": [], "nav": [], "bench": [], "drawdown": [], "holding_count": []}
|
||||
return {"dates": [], "nav": [], "bench": [], "drawdown": [],
|
||||
"holding_count": [], "benchmark_code": None, "index": None}
|
||||
dates = [str(pd.Timestamp(x).date()) for x in df["trade_date"]]
|
||||
return {
|
||||
"dates": [str(pd.Timestamp(x).date()) for x in df["trade_date"]],
|
||||
"dates": dates,
|
||||
"nav": [_num(x) for x in df["nav"]],
|
||||
"bench": [_num(x) for x in df["benchmark_nav"]],
|
||||
"drawdown": [_num(x) for x in df["drawdown"]],
|
||||
"holding_count": [int(x) if x is not None else 0 for x in df["holding_count"]],
|
||||
"benchmark_code": _v(df["benchmark_code"].iloc[-1]),
|
||||
# 可选叠加指数(右轴);不传就是纯净值曲线
|
||||
"index": _index_overlay(index_code, dates, cfg=cfg) if index_code else None,
|
||||
}
|
||||
|
||||
|
||||
def _index_overlay(index_code: str, dates: list[str], *, cfg: Any) -> dict[str, Any]:
|
||||
"""把指数收盘价对齐到净值曲线的日期上,供右轴叠加。
|
||||
|
||||
类目轴上每个类目一个点,序列必须**逐点对齐**(缺的补 None),
|
||||
否则整条指数线会相对净值曲线整体错位。
|
||||
"""
|
||||
df = db.read_sql(
|
||||
"SELECT trade_date, close, index_name FROM hd_index_daily "
|
||||
"WHERE index_code = :c AND trade_date BETWEEN :s AND :e "
|
||||
"ORDER BY trade_date",
|
||||
{"c": index_code, "s": dates[0], "e": dates[-1]}, cfg=cfg,
|
||||
)
|
||||
if df.empty:
|
||||
# 区分「库里没这个指数」(用户传错,应当报错)
|
||||
# 与「这个指数在该区间没有行情」(创业板指对更早的回测),后者只提示。
|
||||
n = int(db.read_sql(
|
||||
"SELECT COUNT(*) AS n FROM hd_index_daily WHERE index_code = :c",
|
||||
{"c": index_code}, cfg=cfg)["n"].iloc[0])
|
||||
if not n:
|
||||
raise HdivError(f"未知指数:{index_code}")
|
||||
return {"code": index_code, "name": index_code, "close": [None] * len(dates),
|
||||
"points": 0, "covered": False}
|
||||
close_by_date = {str(pd.Timestamp(d).date()): _num(c)
|
||||
for d, c in zip(df["trade_date"], df["close"])}
|
||||
close = [close_by_date.get(d) for d in dates]
|
||||
return {
|
||||
"code": index_code,
|
||||
"name": _v(df["index_name"].iloc[0]) or index_code,
|
||||
"close": close,
|
||||
"points": sum(1 for x in close if x is not None),
|
||||
"covered": True,
|
||||
}
|
||||
|
||||
|
||||
@@ -612,7 +855,7 @@ def _reason_text(d: dict[str, Any]) -> str:
|
||||
y = _num(d.get("dividend_yield"))
|
||||
p = _num(d.get("yield_percentile"))
|
||||
if y is not None:
|
||||
parts.append(f"股息率 {y * 100:.2f}%")
|
||||
parts.append(f"股息率 {_fmt().pct(y)}")
|
||||
if p is not None:
|
||||
parts.append(f"历史分位 {p:.1f}%")
|
||||
if d.get("rule"):
|
||||
@@ -652,6 +895,9 @@ _SKIP_LABELS = {
|
||||
"ALREADY_AT_TARGET": "已达目标仓位",
|
||||
"NO_POSITION": "无持仓",
|
||||
"BELOW_MIN_TRADE": "低于最小交易量",
|
||||
# 实时画像闸门剔除(信号类型 REJECT):不是撮合失败,而是「按当日可见
|
||||
# 数据重算画像后判定不值得买」。详情在 reason_json.profile_gate.checks。
|
||||
"PROFILE_GATE": "实时画像未通过,主动放弃买入",
|
||||
}
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user