修复:量价单位 / 未来函数守卫 / 实时画像闸门;行情回补到 2005;手册补全流程
本轮会话的三项正确性改造(均为「不报错、只让结果静默错」的类型):
1) 修复 stock_daily 量价单位前后不一致
- 现象:2015-2019 存 Tushare 原始单位(手/千元),2020 起存(股/元),2019 同日混合;
而流动性阈值按「元」配置 → 早年门槛实际是「日均成交额 ≥ 200 亿元」,
把 2015-2019 的股票池整体清空(实测 2016/2017/2018 各选出 0 只)。
- 修复:写入端 sync/price.py 统一换算;读取端 units.normalize_ohlcv_units
按行判定并幂等换算(price_history / avg_amount 都走它);
审计新增 UNIT-OHLCV 防回归。
- 效果:2016/2017/2018 的股票池变为 7/11/13 只。
2) 未来函数守卫(单次回测)
- 股票池自带 asof:若晚于回测起点即**拒绝执行**(原先静默冻结套用),
与 walk-forward 已有的拒绝理由一致;确需复现加 --allow-lookahead-universe,
偏差写入 unimplemented_json。
3) 新增实时(PIT)个股画像闸门
- profile/pit.py:每个决策日按当时可见数据重算过去 5 年画像,
惰性(仅买入条件已触发的标的)、面板按 asof 缓存、
规则不含财务指标时不查财报表;被剔除时产出 REJECT + 逐规则留痕。
- 指标定义复用 ProfileBuilder._profile_one(与批量画像逐值等价的回归测试)。
- profile/coverage.py:窗口覆盖率(按交易日历的真实开市天数),
策略新增 entry.profile_gate.min_window_coverage(默认 0,不改变既有行为)。
- core/metrics.py:闸门可用指标的唯一定义(配置期即校验,避免写错指标名静默失效)。
4) 行情回补到 2005(使 5/8/10 年窗口真正完整)
- stock_daily / adjust_factor / daily_basic 补到 2005-01-04;
hd_suspend / hd_limit 补到 2010-01-04。
- 5 年窗口覆盖率:2018-05-18 由 67.0% → 99.1%,2016-12-30 由 39.8% → 99.0%;
残差经逐日与 hd_suspend 交叉核实为真实停牌(16/16 命中)。
- 审计 G2/G3 与断点续传原先用固定阈值(2000 / 1500 只),
会把 2005-2009 的正常数据误判为异常 —— 改为按「当年应有上市股票数」成比例判定。
- 节流修正:daily/adj_factor/daily_basic 限频 480 → 170(实测该 token 约 196/min 即被拒)。
5) 自我声明如实化
- 原先「约束未生效」由「过滤后集合为空」判定,会把「这批股票恰好没停牌」
误报成「hd_suspend 无数据」;改为按表级判定。
- 补齐此前静默的「配置承诺但未实现」项:suspended_rule/limit_up_down_rule 的 defer、
cash_mode=reinvest/reinvest_rule、handle_rights_issue、signal_to_execution、
max_volume_pct、liquidity_limit_pct_adv —— 全部写入 unimplemented_json。
6) 手册:新增 §0「全流程操作(选股 → 画像 → 回测)」置于最前
- 逐步说明「命令做了什么、数据从哪来、落了哪些库、有哪些坑」;
含实时画像闸门 9 问 9 答、未来函数守卫表、成交与成本口径、验证 SQL。
- 修正旧 §2.4 漏传 --universe-run(选了池子却没用于回测);
修正两处声称「停牌顺延」「分红再投资」已实现的相反表述。
测试:403 项全部通过(含新增 test_units.py、test_profile_pit.py、
未实现声明诚实性测试、行序无关性回归测试)。
注意:本提交中 docs/*、README.md、src/hdiv/web/service.py 除本轮修改外,
也含此前遗留的未提交改动(无法按文件切分)。
This commit is contained in:
@@ -0,0 +1,138 @@
|
||||
"""量价单位归一化(``stock_daily`` 的历史遗留混用)测试。
|
||||
|
||||
背景:``stock_daily`` 里 2015-01~2019 的行是 Tushare 原始单位(手 / 千元),
|
||||
2020 起沿用既有 qlib 存量(股 / 元),2019 年同日混着两种。
|
||||
按「元」配置的流动性阈值因此把早年低估 1000 倍,会把股票池整体清空。
|
||||
这些测试锁定「读取层必须幂等地归一化」这一契约。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import pandas as pd
|
||||
import pytest
|
||||
|
||||
from hdiv.data.units import (
|
||||
OHLCV_CONVERTED,
|
||||
OHLCV_UNKNOWN,
|
||||
amount_qian_to_yuan,
|
||||
detect_ohlcv_units,
|
||||
normalize_ohlcv_units,
|
||||
ohlcv_unit_ratio,
|
||||
vol_shou_to_shares,
|
||||
)
|
||||
|
||||
|
||||
def _row(symbol: str, close: float, shares: float, **kw) -> dict:
|
||||
"""按「股 / 元」口径构造一行(即换算后的目标形态)。"""
|
||||
return {
|
||||
"symbol": symbol,
|
||||
"close": close,
|
||||
"volume": shares,
|
||||
"amount": shares * close,
|
||||
**kw,
|
||||
}
|
||||
|
||||
|
||||
def _raw_row(symbol: str, close: float, shares: float, **kw) -> dict:
|
||||
"""按 Tushare 原始口径构造一行:成交量为手、成交额为千元。
|
||||
|
||||
``shares`` 是真实股数。手 = 股/100;千元 = (股 × 价)/1000。
|
||||
"""
|
||||
return {
|
||||
"symbol": symbol,
|
||||
"close": close,
|
||||
"volume": shares / 100.0,
|
||||
"amount": shares * close / 1000.0,
|
||||
**kw,
|
||||
}
|
||||
|
||||
|
||||
class TestScalarConversions:
|
||||
def test_vol_shou_to_shares(self) -> None:
|
||||
assert vol_shou_to_shares(pd.Series([100.0])).iloc[0] == 10000.0
|
||||
|
||||
def test_amount_qian_to_yuan(self) -> None:
|
||||
assert amount_qian_to_yuan(pd.Series([1000.0])).iloc[0] == 1_000_000.0
|
||||
|
||||
|
||||
class TestUnitRatio:
|
||||
def test_converted_rows_ratio_is_one(self) -> None:
|
||||
df = pd.DataFrame([_row("000001.SZ", 10.0, 1e6)])
|
||||
assert ohlcv_unit_ratio(df).iloc[0] == pytest.approx(1.0)
|
||||
|
||||
def test_raw_rows_ratio_is_point_one(self) -> None:
|
||||
# 手 / 千元:volume 是股数/100,amount 是元/1000 → 比值 1/10
|
||||
df = pd.DataFrame([_raw_row("000001.SZ", 10.0, 1e6)])
|
||||
assert ohlcv_unit_ratio(df).iloc[0] == pytest.approx(0.1)
|
||||
|
||||
def test_ratio_tolerates_intraday_move(self) -> None:
|
||||
"""VWAP 与收盘价相差 ±10%(涨跌停)时不得误判单位。"""
|
||||
df = pd.DataFrame([
|
||||
_row("A", 10.0, 1e6, amount=1e6 * 11.0), # VWAP 高于收盘 10%
|
||||
_row("B", 10.0, 1e6, amount=1e6 * 9.0), # VWAP 低于收盘 10%
|
||||
])
|
||||
assert list(detect_ohlcv_units(df)) == [OHLCV_CONVERTED, OHLCV_CONVERTED]
|
||||
|
||||
def test_degenerate_rows_are_unknown(self) -> None:
|
||||
df = pd.DataFrame([
|
||||
{"symbol": "A", "close": 0.0, "volume": 100.0, "amount": 1000.0},
|
||||
{"symbol": "B", "close": 10.0, "volume": 0.0, "amount": 1000.0},
|
||||
{"symbol": "C", "close": 10.0, "volume": 100.0, "amount": float("nan")},
|
||||
])
|
||||
assert list(detect_ohlcv_units(df)) == [OHLCV_UNKNOWN] * 3
|
||||
|
||||
|
||||
class TestNormalizeOhlcvUnits:
|
||||
def test_raw_rows_are_converted(self) -> None:
|
||||
raw = pd.DataFrame([{"symbol": "000001.SZ", "close": 9.80,
|
||||
"volume": 417732.0, "amount": 412636.0}])
|
||||
out, diag = normalize_ohlcv_units(raw)
|
||||
assert diag["raw"] == 1 and diag["fixed"] == 1
|
||||
# 41,773,200 股 × 9.8784 ≈ 4.126 亿元
|
||||
assert out.iloc[0]["volume"] == pytest.approx(41_773_200.0)
|
||||
assert out.iloc[0]["amount"] == pytest.approx(412_636_000.0)
|
||||
assert out.iloc[0]["amount"] / out.iloc[0]["volume"] == pytest.approx(9.878, abs=0.01)
|
||||
|
||||
def test_is_idempotent(self) -> None:
|
||||
raw = pd.DataFrame([_raw_row("A", 10.0, 1e6)])
|
||||
once, diag1 = normalize_ohlcv_units(raw)
|
||||
assert diag1["fixed"] == 1
|
||||
twice, diag2 = normalize_ohlcv_units(once)
|
||||
pd.testing.assert_frame_equal(once, twice)
|
||||
assert diag2["raw"] == 0, "已换算的行不得被二次换算"
|
||||
|
||||
def test_mixed_units_within_one_date(self) -> None:
|
||||
"""2019 年同日两种单位并存(实测 3596 行里 237 行已换算)。"""
|
||||
df = pd.DataFrame([
|
||||
_raw_row("RAW", 10.0, 1e6),
|
||||
_row("CONV", 10.0, 1e6),
|
||||
])
|
||||
out, diag = normalize_ohlcv_units(df)
|
||||
assert diag["raw"] == 1 and diag["converted"] == 1
|
||||
for i in out.index:
|
||||
assert out.at[i, "amount"] / (out.at[i, "volume"] * out.at[i, "close"]) == pytest.approx(1.0)
|
||||
|
||||
def test_does_not_mutate_input(self) -> None:
|
||||
raw = pd.DataFrame([_raw_row("A", 10.0, 1e6)])
|
||||
before = raw.copy()
|
||||
normalize_ohlcv_units(raw)
|
||||
pd.testing.assert_frame_equal(raw, before)
|
||||
|
||||
def test_empty_frame(self) -> None:
|
||||
out, diag = normalize_ohlcv_units(pd.DataFrame())
|
||||
assert out.empty and diag["total"] == 0
|
||||
|
||||
def test_missing_column_is_reported_not_guessed(self) -> None:
|
||||
"""缺 close 时无法判定单位 —— 必须原样返回并说明,不得瞎猜。"""
|
||||
df = pd.DataFrame([{"symbol": "A", "volume": 1e4, "amount": 1e5}])
|
||||
out, diag = normalize_ohlcv_units(df)
|
||||
pd.testing.assert_frame_equal(out, df)
|
||||
assert "error" in diag and diag["fixed"] == 0
|
||||
|
||||
def test_custom_column_names(self) -> None:
|
||||
df = pd.DataFrame([_raw_row("A", 10.0, 1e6)].copy())
|
||||
df = df.rename(columns={"volume": "vol", "amount": "amt", "close": "px"})
|
||||
out, diag = normalize_ohlcv_units(df, volume_col="vol", amount_col="amt", close_col="px")
|
||||
assert diag["fixed"] == 1
|
||||
assert out.iloc[0]["vol"] == pytest.approx(1e6)
|
||||
assert out.iloc[0]["amt"] == pytest.approx(1e7)
|
||||
Reference in New Issue
Block a user