功能:每日动态股票池回测(--mode daily)+ 每日增量同步 + PIT 批量取数层
说明:本提交是工作区中此前的未提交工作(在 14ec0c6 之后产生),**非本次会话所写**,
按用户要求整理并推送。已做安全检查(无明文凭据、无大文件、.env/logs/output 仍被忽略),
并完成可执行范围内的测试验证(见「测试」一节)。
## 新增能力
1) `hdiv backtest --mode daily --start <日期>`
- src/hdiv/backtest/daily.py:两趟式(先逐日选股,再复用既有引擎模拟)
- 每个交易日按当日可见数据重建股票池(PIT),每个交易日判断买卖点
- `pool_exit_action`:hold(只减不加、不因掉出池子而清仓)/ sell(掉出即清仓)
- `profile_on_trade`:买卖决策发生时计算并留痕个股画像,**不区分是否在当日池内**
(卖出/减仓同样留痕,否则「为什么卖」缺证据)
- 与 walkforward 的分工:daily 是一条连续路径的推演,不是过拟合检验;
因此不使用训练段、不冻结分布,阈值口径一律 rolling
- 拒绝 `--universe-run`(daily 的定义就是逐日重筛,冻结池与之矛盾)
2) PIT 批量取数层 src/hdiv/universe/pit.py
- PitRepo 继承 Repo,**只重写取数**(按区块批量预载 + 逐日内存切片),
派生逻辑(最新一期财报合并、单位归一化、支付率口径等)一行不重写
—— 以保证与逐日单点查询**结果等价**
- 候选集预剪枝:用「不可能通过」的边界条件提前排除,文档论证为精确等价而非近似
- src/hdiv/universe/daily.py:每日动态筛选器(仍然调用既有 selector 与四个 Filter)
3) 每日增量同步 `hdiv sync daily`
- src/hdiv/data/sync/daily.py:只抓「库里还没有的那几天」,
按「当日股票数 ≥ 当年规模阈值」判定缺口,不重拉历史、不覆盖既有行;
支持 `--dry-run` 先看待抓清单
- deploy/daily-sync.sh、deploy/install-sync-schedule.sh、
deploy/com.hddiv.sync.plist.example(launchd 每天 17:00)
- 新表 hd_daily_universe(逐日入选成员留痕)+ sql/hd_daily_universe.sql + schema.py
(该表已存在于库中,`ddl plan` 返回 0 个待执行动作)
4) Web 与文档
- 前端支持 daily 模式记录下钻(web/app.js、web/app.css、web/index.html、
web/favicon.svg)
- README / docs/user-guide.md / docs/implementation-status.md 同步更新:
三种回测模式的取舍、daily 的成本说明(6.7 年约 1.5 小时)与调优手段
## 测试
tests/ 共 500 项(新增 tests/test_daily.py 43 项、tests/test_sync_daily.py 36 项)。
已验证通过:
- 排除上述两个新文件的 **421 项:全部通过(pytest 退出码 0)**
- 两个新文件的**非 DB 单元测试 60 项:全部通过**
未能在合理时间内跑完:
- 两个新文件中 **19 项 DB 标记的重型测试**。实测瓶颈是一条**无界全表扫描**:
`SELECT ... FROM hd_cashflow WHERE ann_date <= :asof ORDER BY symbol, end_date, ann_date`
(31 万行,无 symbol/报告期下限)。全量套件跑到 161 项时已耗时 20 分钟、
0 失败,按该速率预计需 3 小时以上,因此改为分档验证。
- 旁证:库中存在 3 次成功的 daily 端到端运行(2026-10-05 10:05 / 10:32 / 11:03,
区间 2024-03-01~03-15),说明该路径可正常完成。
## 已知待改进
- 上述 `hd_cashflow`(及同类「按 ann_date 上界取全历史」)的查询缺
symbol / 报告期下限,是 daily 模式的主要性能瓶颈,建议下一轮优化。
This commit is contained in:
@@ -43,12 +43,16 @@ FRONTEND_CALLS: list[tuple[str, str]] = [
|
||||
# 净值曲线右轴可叠加的基准指数
|
||||
("GET", "/api/indices"),
|
||||
("GET", "/api/backtests/abc123/signals"),
|
||||
# 每日动态股票池(--mode daily):时间线 / 某日成员
|
||||
("GET", "/api/backtests/abc123/daily-universe"),
|
||||
("PATCH", "/api/backtests/abc123"),
|
||||
# 回测内分析:任意日持仓 + 个股买卖点
|
||||
("GET", "/api/backtests/abc123/portfolio"),
|
||||
("GET", "/api/backtests/abc123/position-dates"),
|
||||
("GET", "/api/backtests/abc123/stocks"),
|
||||
("GET", "/api/backtests/abc123/stocks/600519.SH"),
|
||||
# 已清仓了结清单(含清仓后至今涨跌)
|
||||
("GET", "/api/backtests/abc123/closed-positions"),
|
||||
# Walk-forward 样本外
|
||||
("GET", "/api/walkforwards"),
|
||||
("GET", "/api/walkforwards/abc123"),
|
||||
@@ -179,6 +183,78 @@ def test_frontend_uses_hash_routing_only() -> None:
|
||||
assert "pushState" not in js
|
||||
|
||||
|
||||
def test_router_sentinel_is_not_the_home_path() -> None:
|
||||
"""路由的「已渲染路径」哨兵不能是空串 —— 首页路径本身就是 ''。
|
||||
|
||||
历史 bug:``let currentPath = ''`` 且所有强制重渲染都写 ``currentPath = ''``。
|
||||
首页(hash 为空)解析出的 path 恰好也是 '',于是 render() 一进门就命中
|
||||
``path === currentPath`` 提前返回:index.html 里那句「加载中…」永远不被替换,
|
||||
概览页整页打不开(其他页面因为有非空路径,反而正常)。哨兵改用 null 后,
|
||||
'' 才能被当作一个正常的、需要渲染的路径。
|
||||
"""
|
||||
js = (project_root() / "web" / "app.js").read_text(encoding="utf-8")
|
||||
# 只看赋值(排除 === 比较)
|
||||
assigns = [a.strip() for a in re.findall(r"currentPath\s*=(?!=)\s*([^;\n]+)", js)]
|
||||
assert assigns, "未在 app.js 中找到 currentPath 赋值,解析逻辑需更新"
|
||||
bad = [a for a in assigns if a in {"''", '""'}]
|
||||
assert not bad, f"currentPath 不能用空串作哨兵(与首页路径 '' 冲突):{bad}"
|
||||
assert "let currentPath = null" in js
|
||||
|
||||
|
||||
def test_profile_gate_is_exposed_for_display() -> None:
|
||||
"""画像闸门是买入判据的一部分,必须能在回测页看到。
|
||||
|
||||
看不到就会出现「股息率分位到了却没买」无从解释的情况 ——
|
||||
闸门规则是**第二道**买入条件,和 run 一起要能复现。
|
||||
"""
|
||||
from hdiv.web.service import describe_strategy
|
||||
|
||||
cfg = {
|
||||
"strategy": {"id": "S", "name": "n", "version": "1", "status": "DRAFT",
|
||||
"description": ""},
|
||||
"entry": {
|
||||
"yield_percentile": 75,
|
||||
"profile_gate": {
|
||||
"enabled": True, "window_years": 5, "on_unverifiable": "reject",
|
||||
"min_window_coverage": 0.0,
|
||||
"rules": [
|
||||
{"metric": "payout_ratio", "stat": "current_value",
|
||||
"op": "<=", "value": 1.0},
|
||||
# 没写 stat:应默认当日值,且不能把规则丢掉
|
||||
{"metric": "roe_avg", "op": ">=", "value": 0.08},
|
||||
],
|
||||
},
|
||||
},
|
||||
}
|
||||
g = describe_strategy(cfg)["profile_gate"]
|
||||
assert g["enabled"] is True and g["window_years"] == 5.0
|
||||
assert g["on_unverifiable"] == "reject"
|
||||
assert [r["metric"] for r in g["rules"]] == ["payout_ratio", "roe_avg"]
|
||||
assert g["rules"][0]["op"] == "<=" and g["rules"][0]["value"] == 1.0
|
||||
assert g["rules"][1]["stat"] == "current_value"
|
||||
json.dumps(g, ensure_ascii=False, allow_nan=False)
|
||||
|
||||
# 老配置没有这一段 → None,前端据此不显示卡片
|
||||
assert describe_strategy({"strategy": {}, "entry": {}})["profile_gate"] is None
|
||||
# 脏数据不能把整页带崩,也不能造出假规则
|
||||
dirty = describe_strategy({"strategy": {},
|
||||
"entry": {"profile_gate": {"enabled": True,
|
||||
"rules": [None, {}, "x"]}}})
|
||||
assert dirty["profile_gate"]["rules"] == []
|
||||
|
||||
|
||||
def test_profile_gate_card_is_wired_into_backtest_page() -> None:
|
||||
"""回测页必须真的把画像闸门渲染出来(接口有字段≠页面显示)。"""
|
||||
js = (project_root() / "web" / "app.js").read_text(encoding="utf-8")
|
||||
assert "profileGateCard" in js, "缺少画像闸门卡片渲染函数"
|
||||
assert "个股画像筛选条件" in js, "缺少画像闸门卡片标题"
|
||||
assert "profile_gate" in js, "未把接口字段接到卡片上"
|
||||
# 卡片要挂在「回测条件」之后
|
||||
cond = js.index(">回测条件<")
|
||||
gate = js.index("profileGateCard(b.strategy.profile_gate)")
|
||||
assert cond < gate, "画像闸门卡片必须在「回测条件」之后"
|
||||
|
||||
|
||||
def test_frontend_escapes_html() -> None:
|
||||
"""用户可输入记录名称/备注,必须转义以避免 XSS。"""
|
||||
js = (project_root() / "web" / "app.js").read_text(encoding="utf-8")
|
||||
@@ -383,6 +459,139 @@ def test_equity_index_overlay_is_date_aligned() -> None:
|
||||
service.get_backtest_equity(rid, index_code="999999.XX")
|
||||
|
||||
|
||||
@requires_db
|
||||
def test_stock_detail_default_range_reaches_latest_data() -> None:
|
||||
"""回归:默认区间要到**该股最新行情**,而不是持仓结束(卖出)当天。
|
||||
|
||||
老实现默认用「持仓区间」,卖出之后曲线就断了,
|
||||
「卖飞了没有」这个最该回答的问题在图上无从回答。
|
||||
这里特意挑「已清仓、且清仓日之后还有行情」的样本 —— 正是老实现会断线的场景。
|
||||
"""
|
||||
from hdiv.core.config import load_config
|
||||
from hdiv.data import db
|
||||
from hdiv.web import analysis
|
||||
|
||||
cfg = load_config("datasource")
|
||||
df = db.read_sql(
|
||||
"SELECT p.run_id, p.symbol, p.hold_start, p.hold_end, d.avail_end "
|
||||
"FROM (SELECT run_id, symbol, MIN(trade_date) AS hold_start, "
|
||||
" MAX(trade_date) AS hold_end "
|
||||
" FROM hd_backtest_position GROUP BY run_id, symbol) p "
|
||||
"JOIN (SELECT symbol, MAX(trade_date) AS avail_end "
|
||||
" FROM daily_basic GROUP BY symbol) d ON d.symbol = p.symbol "
|
||||
"WHERE p.hold_end < d.avail_end "
|
||||
"ORDER BY p.hold_end LIMIT 1",
|
||||
cfg=cfg,
|
||||
)
|
||||
if df.empty:
|
||||
pytest.skip("库里没有「已清仓且之后仍有行情」的样本")
|
||||
row = df.iloc[0]
|
||||
rid, sym = str(row["run_id"]), str(row["symbol"])
|
||||
hold_start, hold_end, avail_end = (str(row["hold_start"]), str(row["hold_end"]),
|
||||
str(row["avail_end"]))
|
||||
|
||||
r = analysis.stock_detail(rid, sym, series=["close"])["range"]
|
||||
assert r["available_start"] and r["available_end"], "应返回该股行情边界供日期选择器用"
|
||||
assert r["end"] == avail_end, \
|
||||
f"默认区间止于 {r['end']},而行情已到 {avail_end}(又回到「卖出即断线」)"
|
||||
assert r["start"] <= hold_start, "默认起点不应晚于持仓起点(判据数据要在图上)"
|
||||
assert r["end"] > hold_end, f"默认区间不应停在清仓日 {hold_end}"
|
||||
|
||||
|
||||
@requires_db
|
||||
def test_stock_detail_default_range_includes_judgement_lookback() -> None:
|
||||
"""默认区间要含**首笔成交之前**的判据数据。
|
||||
|
||||
买入依据是股息率的历史分位(滚动窗口,config: percentile_reference
|
||||
.lookback_years = 5 年);只画持仓期等于把「当时凭什么买」的判据裁掉了。
|
||||
"""
|
||||
import datetime as _dt
|
||||
|
||||
from hdiv.web import analysis
|
||||
|
||||
rid = _sample_backtest_run()
|
||||
if not rid:
|
||||
pytest.skip("没有可用的回测")
|
||||
stocks = [s for s in analysis.run_stocks(rid) if (s.get("trade_count") or 0) > 0]
|
||||
if not stocks:
|
||||
pytest.skip("该回测没有成交")
|
||||
sym = stocks[0]["symbol"]
|
||||
|
||||
d = analysis.stock_detail(rid, sym, series=["close"])
|
||||
r = d["range"]
|
||||
first = min(t["execution_date"] for t in d["trades"])
|
||||
need = _dt.date.fromisoformat(first) - _dt.timedelta(days=int(365.25 * 5))
|
||||
assert r["start"] <= need.isoformat(), \
|
||||
f"默认起点 {r['start']} 未覆盖首笔成交({first})之前 5 年的判据数据"
|
||||
# 但不该早于该股行情本身(否则日期选择器会给出选不到的日期)
|
||||
assert r["start"] >= r["available_start"]
|
||||
|
||||
|
||||
@requires_db
|
||||
def test_downsampling_keeps_trade_dates() -> None:
|
||||
"""回归:降采样不能把成交日丢掉。
|
||||
|
||||
前端按**日期**把买卖点落到横轴上,横轴里没有那一天,
|
||||
这笔成交就会从图上消失(还会被前端误报成「不在所选区间内」)。
|
||||
实测:默认区间放宽到「5 年判据 + 至今」后,13 只降采样股票里有 7 只会丢成交日。
|
||||
"""
|
||||
from hdiv.web import analysis
|
||||
|
||||
rid = _sample_backtest_run()
|
||||
if not rid:
|
||||
pytest.skip("没有可用的回测")
|
||||
checked = 0
|
||||
for s in analysis.run_stocks(rid):
|
||||
d = analysis.stock_detail(rid, s["symbol"], series=["close"])
|
||||
if not d["range"]["downsampled"]:
|
||||
continue
|
||||
checked += 1
|
||||
axis = set(d["dates"])
|
||||
missing = [t["execution_date"] for t in d["trades"]
|
||||
if t["execution_date"] not in axis]
|
||||
assert not missing, \
|
||||
f"{s['symbol']} 降采样后丢了成交日 {missing},图上会少标这几笔"
|
||||
assert d["dates"] == sorted(d["dates"]), "横轴仍须按时间升序"
|
||||
if not checked:
|
||||
pytest.skip("该回测没有触发降采样的个股")
|
||||
|
||||
|
||||
@requires_db
|
||||
def test_stock_detail_date_params_and_validation() -> None:
|
||||
"""区间参数:能收窄、非法输入报可读错误、区间外成交仍要返回。"""
|
||||
import datetime as _dt
|
||||
|
||||
from hdiv.core.errors import HdivError
|
||||
from hdiv.web import analysis
|
||||
|
||||
rid = _sample_backtest_run()
|
||||
if not rid:
|
||||
pytest.skip("没有可用的回测")
|
||||
stocks = analysis.run_stocks(rid)
|
||||
if not stocks:
|
||||
pytest.skip("该回测没有持仓股票")
|
||||
sym = stocks[0]["symbol"]
|
||||
|
||||
base = analysis.stock_detail(rid, sym, series=["close"])
|
||||
a = _dt.date.fromisoformat(base["range"]["start"])
|
||||
s, e = a.isoformat(), (a + _dt.timedelta(days=180)).isoformat()
|
||||
|
||||
d = analysis.stock_detail(rid, sym, start=s, end=e, series=["close"])
|
||||
assert d["range"]["requested_start"] == s and d["range"]["requested_end"] == e
|
||||
assert s <= d["range"]["start"] and d["range"]["end"] <= e
|
||||
assert d["range"]["points"] < base["range"]["points"], "收窄区间应真的少取数据"
|
||||
assert d["range"]["default_end"] == base["range"]["default_end"], \
|
||||
"default_* 应是「重置」用的缺省区间,不随本次请求变化"
|
||||
# 区间外的成交仍要返回:前端靠它提示「有 N 笔不在所选区间内」
|
||||
assert len(d["trades"]) == len(base["trades"])
|
||||
|
||||
with pytest.raises(HdivError):
|
||||
analysis.stock_detail(rid, sym, start=e, end=s, series=["close"])
|
||||
for bad in ("2024-13-45", "not-a-date"):
|
||||
with pytest.raises(HdivError):
|
||||
analysis.stock_detail(rid, sym, start=bad, series=["close"])
|
||||
|
||||
|
||||
@requires_db
|
||||
def test_reason_text_is_human_readable() -> None:
|
||||
"""成交理由必须渲染成人话,而不是丢一坨 JSON 给前端。"""
|
||||
@@ -632,6 +841,22 @@ def test_site_build_does_not_clobber_spa() -> None:
|
||||
assert "app/app.js" in html
|
||||
|
||||
|
||||
def test_published_site_is_world_readable() -> None:
|
||||
"""发布产物必须 world-readable:nginx worker 以 nobody 运行,不是文件属主。
|
||||
|
||||
``shutil.copy2`` 会保留源文件权限,所以一个 umask 077 存下来的 600 文件
|
||||
会让线上 CSS/JS 直接 403(HTML 打得开、页面裸奔)。发布时统一收敛权限。
|
||||
"""
|
||||
from hdiv.web import site
|
||||
|
||||
site.sync_frontend(verbose=False)
|
||||
out = project_root() / "output"
|
||||
unreadable = [p for p in out.rglob("*") if p.is_file() and not p.stat().st_mode & 0o044]
|
||||
assert not unreadable, f"这些发布文件 nginx(nobody)读不到:{unreadable[:5]}"
|
||||
untraversable = [p for p in out.rglob("*") if p.is_dir() and not p.stat().st_mode & 0o011]
|
||||
assert not untraversable, f"这些目录 nginx(nobody)进不去:{untraversable[:5]}"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 回测内分析:任意日持仓 + 个股买卖点
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -651,6 +876,80 @@ def _sample_backtest_run() -> str | None:
|
||||
return None if df.empty else str(df["run_id"].iloc[0])
|
||||
|
||||
|
||||
@requires_db
|
||||
def test_closed_positions_definition_and_math() -> None:
|
||||
"""已清仓清单:定义(期末不再持有)+ 口径(收益率、清仓后涨跌)都要对得上。
|
||||
|
||||
「已清仓」若按成交净额判断会漏掉「卖了又买回、期末仍持有」的票,
|
||||
这里同时用两套口径交叉验证,并要求金额/盈亏与成交表逐笔汇总一致。
|
||||
"""
|
||||
from hdiv.core.config import load_config
|
||||
from hdiv.data import db
|
||||
from hdiv.core.errors import HdivError
|
||||
from hdiv.web import analysis
|
||||
|
||||
cfg = load_config("datasource")
|
||||
# 找一只有已清仓个股的回测(期末持仓数 < 曾持有数)
|
||||
df = db.read_sql(
|
||||
"SELECT run_id FROM hd_backtest_position GROUP BY run_id "
|
||||
"HAVING COUNT(DISTINCT symbol) > "
|
||||
" (SELECT COUNT(DISTINCT symbol) FROM hd_backtest_position p2 "
|
||||
" WHERE p2.run_id = hd_backtest_position.run_id "
|
||||
" AND p2.trade_date = (SELECT MAX(trade_date) FROM hd_backtest_position p3 "
|
||||
" WHERE p3.run_id = hd_backtest_position.run_id)) "
|
||||
"ORDER BY COUNT(DISTINCT symbol) DESC LIMIT 1",
|
||||
cfg=cfg,
|
||||
)
|
||||
if df.empty:
|
||||
pytest.skip("没有含已清仓个股的回测")
|
||||
rid = str(df["run_id"].iloc[0])
|
||||
|
||||
d = analysis.closed_positions(rid)
|
||||
items = d["items"]
|
||||
assert items, "该回测应当有已清仓个股"
|
||||
json.dumps(d, ensure_ascii=False, allow_nan=False) # NaN 不能漏到前端
|
||||
|
||||
last_day = str(db.read_sql(
|
||||
"SELECT MAX(trade_date) AS d FROM hd_backtest_position WHERE run_id = :r",
|
||||
{"r": rid}, cfg=cfg)["d"].iloc[0])
|
||||
still = set(db.read_sql(
|
||||
"SELECT DISTINCT symbol FROM hd_backtest_position "
|
||||
"WHERE run_id = :r AND trade_date = :d", {"r": rid, "d": last_day}, cfg=cfg)["symbol"])
|
||||
|
||||
tr = db.read_sql(
|
||||
"SELECT symbol, side, quantity, amount, realized_pnl, execution_date "
|
||||
"FROM hd_backtest_trade WHERE run_id = :r", {"r": rid}, cfg=cfg)
|
||||
for x in items:
|
||||
assert x["symbol"] not in still, f"{x['symbol']} 期末仍持有,不该出现在已清仓清单"
|
||||
mine = tr[tr["symbol"] == x["symbol"]]
|
||||
buys = mine[mine["side"] == "BUY"]
|
||||
sells = mine[mine["side"] == "SELL"]
|
||||
assert len(sells) > 0, "已清仓必然有卖出成交"
|
||||
# 刻意**不**校验「买入股数 == 卖出股数」:送股/转增会让持仓股数凭空增加
|
||||
# (实测 600188.SH 在 93fb7456 里买入 6000 股、卖出 11700 股)。
|
||||
# 所以「已清仓」只能以持仓表为准,不能用成交净额反推。
|
||||
assert x["last_sell"] == str(sells["execution_date"].max())
|
||||
assert abs((x["realized_pnl"] or 0) - float(sells["realized_pnl"].sum())) < 1e-6
|
||||
assert abs((x["buy_amount"] or 0) - float(buys["amount"].sum())) < 1e-6
|
||||
assert x["first_hold"] and x["last_hold"] and x["hold_days"] > 0
|
||||
if x["return_pct"] is not None: # 已清仓 ⇒ 收益率 = 已实现盈亏 / 买入金额
|
||||
assert abs(x["return_pct"] - x["realized_pnl"] / x["buy_amount"]) < 1e-9
|
||||
if x["since_sell_pct"] is not None: # 清仓后涨跌以清仓日收盘为基准
|
||||
assert abs(x["since_sell_pct"] -
|
||||
(x["close_latest"] / x["close_at_sell"] - 1.0)) < 1e-9
|
||||
|
||||
s = d["summary"]
|
||||
assert s["count"] == len(items)
|
||||
assert abs(s["realized_pnl"] - sum(x["realized_pnl"] or 0 for x in items)) < 1e-6
|
||||
assert s["since_sell_up"] + s["since_sell_down"] <= s["count"]
|
||||
# 明细按清仓日倒序(最近清仓的排在最前)
|
||||
dates = [x["last_sell"] or "" for x in items]
|
||||
assert dates == sorted(dates, reverse=True)
|
||||
|
||||
with pytest.raises(HdivError):
|
||||
analysis.closed_positions("不存在的runid")
|
||||
|
||||
|
||||
@requires_db
|
||||
def test_position_dates_is_compact_by_default() -> None:
|
||||
"""默认只返回日期字符串:带全字段会让响应从约 30KB 涨到 460KB。"""
|
||||
|
||||
Reference in New Issue
Block a user