说明:本提交是工作区中此前的未提交工作(在 cf6d4d2 之后产生),**非本次会话所写**,
按用户要求**不跑测试、直接记录变更并推送**。
已完成推送前的基础安全检查:无明文凭据、无大文件、`.env`/`logs/`/`output/` 仍被忽略。
测试状态:**本次未执行测试套件**。
## 一、TTM 股息率的两处残余缺陷 + 卖出复核
起因:用户报告 `600690.SH` 在 2026-07-30 触发清仓、07-31 开盘卖出,实际不该卖。
### 缺陷一:同一除权日的多条「实施」记录被逐行累加
- 成因:`hd_dividend` 写入侧刻意保留全量公告记录,去重键含 `ann_date`,
同一笔分红会有多条「实施」记录落在**同一除权日**;查询侧逐行累加即重复计入。
- 规模:5724 只有现金分红的股票中 **953 只**存在同除权日重复(多出 1313 行)。
- 效果:`600690.SH` 的 `ttm_dps` 长期虚高约一倍(2.46848 vs 真实 1.23424),
窗口到期时又必然回落,把假象放大成一次 −78% 的塌陷。
- 修法:新增 `factor.dividend_yield.dedupe_dividend_events()`,按 `(symbol, ex_date)`
聚合成**一笔经济事件**(金额/股数逐字段取最大 → 收敛「分项 + 合计」;
日期取最晚 → PIT 保守)。三处入口统一调用:`ttm_dps_series`、
`Repo.dividend_events`、`universe/filters/dividend.py`。
### 缺陷二:只看相邻间隔,漏掉「年度 → 中期 → 下一年度」的跳法
- 成因:7.6 的「按后继接管」只看相邻两次除权的间隔。实测 `600690.SH`:
FY2024 年度 2025-07-25、FY2025 中期 2025-11-07、FY2025 年度 2026-08-21。
105 天的间隔使前两笔被判为「年内多次分红」而互不取代,392 天又超过 `365+45`
→ **2026-07-25~08-21 出现 28 天空窗**,可见现金只剩 0.26920。
- 修法:`ttm_dps_series` 的覆盖窗口由「按相邻间隔」升级为「**按财年 `end_date`**」:
① 后继接管(保留 7.6 行为,阈值 `ttm_days - grace_days` = 320 天);
② **跨财年补位**:每个财年最后一笔 → 下一财年最后一笔入场,上限 `365 + grace`;
③ **末笔宽限兜底**:无后继时覆盖 `365 + grace`(真停发仍如实归零)。
- 验收(作者实测):600690 在 2026-07-27 的 `ttm_dps` 由 0.53840 变为 **1.23424**,
股息率 5.30%、历史分位 92.98%,**不再触发 P25 清仓**。
### 缺陷三(设计缺口):卖出只认「已除权的现金」,不认「已公告的分红」
- 成因:FY2025 年度分红 0.89151 的**实施公告日是 2026-06-25**,除权日 2026-08-21。
TTM 现金口径看不到它 → 「股息率处于历史低位」在字面上为真,
实际描述的是**现金流时点**而非分红能力恶化。
- 修法:新增 `entry/exit.confirm`(`enabled` / `min_ratio` / `announce_lookback_days`):
若「已公告未除权」的分红说明股息率本应更高,且
`TTM ÷ (TTM + 已公告未除权) < min_ratio`,则判定**未确认**:
保持仓位并记录 `EXIT_UNCONFIRMED`(不进成交流水)。真降息不会命中。
## 二、公司行为的三处静默错误(分红/送转/配股口径)
### 问题一:纯送转被整行丢弃(凭空亏损)
- `_apply_dividends` 在算送股**之前**就按 `cash_div_tax <= 0` 整行 `continue`,
于是「10 送 10」这类**无现金分红**的送转完全不调股数 ——
而价格是不复权价、除权日照常腰斩 → 记出一笔不存在的亏损。
- 规模:全库「实施且 `stk_div > 0`」13,038 行,其中**纯送转 3,268 行**;
高股息池成员在 2015-2026 区间内 **824 笔**(如 `000793.SZ` 每 10 股转增 12 股,
单笔约 −54% 的该持仓市值)。
- 修法:现金与送转**各自独立判断**,只有「既无现金也无送转」才跳过;
并把 `stk_bo_rate`/`stk_co_rate` 写入分红台账留痕。
### 问题二:同一除权日的重复记录被重复入账
- 全库 **1401 组**同 `(symbol, ex_date)` 的多条实施记录(1240 组字段相同;
96 组报告期不同、122 组金额不同)。实测 `002352.SZ 2024-11-07` 同时有
0.4 / 1.0 / 1.4 三条,而 1.4 = 0.4 + 1.0 是合计口径 → 逐行累加会放大两三倍。
- 修法:复用 `dedupe_dividend_events`(与缺陷一同一个函数)。
### 问题三:分红再投资的声明与行为不一致
- 引擎实际行为一直是「分红现金回到与初始资金同一个 `cash` 变量,
下次调仓按目标权重再配置」= `reinvest` + `portfolio_rebalance`;
但 `backtest.yml` 写的是 `same_stock_next_open`,于是每次 run 都声明
「未实现分红再投资规则,分红留存为现金」,让人误以为分红不可再投资。
- 修法:配置改为已实现组合 `cash_mode: reinvest` + `reinvest_rule: portfolio_rebalance`;
声明逻辑抽成 `dividend_handling_notes()`,**逐档取值都有单测**对应
(`hold`/`cash_out`/`same_stock_next_open`/`handle_stock_dividend=false`/配股
才声明未实现)。顺带接线一直是**死字段**的 `dividend.apply_dividend_tax`。
## 三、新增:回测层面的排除行业清单(黑名单)
- 位置与语义:`config/backtest.yml: universe_exclusions.industries` ——
「**这次回测**特意不要哪些行业」(研究口径),
与 `config/universe.yml`(策略选股定义)**叠加取并集**,只做减法。
三种回测模式(single / walkforward / daily)一律生效。
- 最大的坑:数据库 `stock.industry` 里**没有「房地产业」**,它被拆成四个名字,
写「房地产」或「房地产业」**一只都排除不掉**:
`全国地产` 26 只 + `区域地产` 43 只 + `房产服务` 13 只 + `园区开发` 14 只 = **96 只**(1.6%)。
因此 `MarketFilter` 首次求值时拿名单与表内实际取值核对,
**写错名字直接抛 `ConfigError`**(并按字符重合度提示最接近的真实取值)。
- 接线:三个入口都走生效后的配置;并修掉一处缓存陷阱(配置变更后缓存未失效)。
- 新增测试锁定它。
## 四、其它
- `src/hdiv/core/config.py`:新增配置模型(排除行业、卖出复核等,+102 行)
- `src/hdiv/data/repo.py`(+45)、`src/hdiv/backtest/engine.py`(+121)、
`backtest/daily.py`、`backtest/walk_forward.py`、`web/service.py`、
`report/universe_report.py` 相应接线
- 测试:新增 `tests/test_dividend_fiscal_year.py`;扩充
`test_backtest.py` / `test_config.py` / `test_daily.py` /
`test_dividend_smoothing.py` / `test_universe.py`
- `tools/diag_dividend_artifact.py`:诊断脚本与上述修复对齐
- 文档:`docs/implementation-status.md` 新增 §7.6b / §7.7 / §11;
`docs/user-guide.md` 新增排除行业清单说明
## 待验证
本次按要求**未执行测试**。上述「实测/验收」数字均引自文档中作者自己的记录,
非本次会话验证结果。建议合入后跑一次全量测试(注意:daily 的 DB 标记测试
因 `hd_cashflow` 无界扫描仍然很慢)。
742 lines
30 KiB
Python
742 lines
30 KiB
Python
"""单位换算与筛选滤网测试。
|
||
|
||
这里覆盖的都是**开发过程中真实踩到过**的坑,每个测试对应一次静默错误:
|
||
|
||
1. ``total_mv`` 是万元而阈值配的是元 → 筛选结果为空(不报错);
|
||
2. 拿季报 ROE(年初至今累计)去比「5 年年均 ROE」→ 好公司全被误杀;
|
||
3. 银行负债率天然 90%+ → 整个金融板块被误杀;
|
||
4. 分红除权晚于 asof → 稳定分红公司被误判为「连续分红 0 年」。
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
from datetime import date
|
||
|
||
import pandas as pd
|
||
import pytest
|
||
|
||
from hdiv.data.units import (
|
||
normalize_financial_panel,
|
||
normalize_market_panel,
|
||
pct_to_ratio,
|
||
verify_market_units,
|
||
vol_shou_to_shares,
|
||
wan_to_shares,
|
||
wan_to_yuan,
|
||
)
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# 单位换算
|
||
# ---------------------------------------------------------------------------
|
||
|
||
|
||
def test_wan_to_yuan() -> None:
|
||
assert wan_to_yuan(pd.Series([1.0, 100.0])).tolist() == [1e4, 1e6]
|
||
|
||
|
||
def test_wan_to_shares() -> None:
|
||
assert wan_to_shares(pd.Series([125619.78])).iloc[0] == pytest.approx(1.2561978e9)
|
||
|
||
|
||
def test_pct_to_ratio() -> None:
|
||
assert pct_to_ratio(pd.Series([5.08])).iloc[0] == pytest.approx(0.0508)
|
||
|
||
|
||
def test_vol_shou_to_shares() -> None:
|
||
assert vol_shou_to_shares(pd.Series([100])).iloc[0] == 10000
|
||
|
||
|
||
def test_normalize_market_panel_units() -> None:
|
||
"""茅台 2024-06-28 实测值:总市值 1.84 万亿元,总股本 12.56 亿股。"""
|
||
df = pd.DataFrame(
|
||
{
|
||
"symbol": ["600519.SH"],
|
||
"close": [1467.39],
|
||
"total_share": [125619.78], # 万股
|
||
"total_mv": [184333208.97], # 万元
|
||
"circ_mv": [184333208.97],
|
||
"dv_ttm": [5.1720], # 百分数
|
||
"turnover_rate": [0.25],
|
||
}
|
||
)
|
||
out = normalize_market_panel(df)
|
||
assert out["total_mv"].iloc[0] == pytest.approx(1.8433320897e12, rel=1e-6)
|
||
assert out["total_share"].iloc[0] == pytest.approx(1.2561978e9, rel=1e-9)
|
||
assert out["dv_ttm"].iloc[0] == pytest.approx(0.051720)
|
||
assert out["_units"].iloc[0] == "yuan/shares/ratio"
|
||
|
||
|
||
def test_normalize_market_panel_does_not_mutate_input() -> None:
|
||
df = pd.DataFrame({"total_mv": [100.0], "total_share": [10.0]})
|
||
before = df.copy()
|
||
normalize_market_panel(df)
|
||
pd.testing.assert_frame_equal(df, before)
|
||
|
||
|
||
def test_normalize_financial_panel_converts_pct() -> None:
|
||
df = pd.DataFrame({"roe": [15.13], "debt_to_assets": [90.23], "total_revenue": [1.78e11]})
|
||
out = normalize_financial_panel(df)
|
||
assert out["roe"].iloc[0] == pytest.approx(0.1513)
|
||
assert out["debt_to_assets"].iloc[0] == pytest.approx(0.9023)
|
||
assert out["total_revenue"].iloc[0] == 1.78e11, "金额列本就以元计,不得被换算"
|
||
|
||
|
||
def test_verify_market_units_detects_correct_data() -> None:
|
||
df = pd.DataFrame(
|
||
{
|
||
"close": [10.0, 20.0],
|
||
"total_share": [1e9, 5e8],
|
||
"total_mv": [1e10, 1e10],
|
||
}
|
||
)
|
||
v = verify_market_units(df)
|
||
assert v["checked"] == 2 and v["bad"] == 0
|
||
assert v["median_ratio"] == pytest.approx(1.0)
|
||
|
||
|
||
def test_identity_alone_cannot_detect_wan_vs_yuan() -> None:
|
||
"""恒等式对「整体万元/元混淆」无效 —— 因为 元/股 × 万股 = 万元。
|
||
|
||
这是一个容易误以为「有了恒等式检查就安全」的陷阱,必须显式记录:
|
||
原始单位下比值仍然恰好是 1。
|
||
"""
|
||
raw = pd.DataFrame(
|
||
{"close": [1467.39], "total_share": [125619.78], "total_mv": [184333208.97]}
|
||
)
|
||
v = verify_market_units(raw)
|
||
assert v["identity_ok"] is True, "恒等式在原始单位下同样成立"
|
||
assert v["unit_ok"] is False, "但绝对量级检查必须发现单位错误"
|
||
assert v["detected_unit"] == "wan"
|
||
|
||
|
||
def test_verify_market_units_accepts_normalized_yuan() -> None:
|
||
norm = normalize_market_panel(
|
||
pd.DataFrame(
|
||
{"close": [1467.39], "total_share": [125619.78], "total_mv": [184333208.97]}
|
||
)
|
||
)
|
||
v = verify_market_units(norm)
|
||
assert v["unit_ok"] is True and v["identity_ok"] is True
|
||
assert v["detected_unit"] == "yuan"
|
||
|
||
|
||
def test_verify_market_units_detects_partial_conversion() -> None:
|
||
"""只换算一个字段(常见疏漏)也必须被抓出来。"""
|
||
df = pd.DataFrame(
|
||
{"close": [10.0], "total_share": [1e9], "total_mv": [1e10 / 1e4]} # 市值漏换算
|
||
)
|
||
v = verify_market_units(df)
|
||
assert v["bad"] == 1
|
||
|
||
|
||
def test_verify_market_units_empty() -> None:
|
||
assert verify_market_units(pd.DataFrame())["checked"] == 0
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# 滤网:共用夹具
|
||
#
|
||
# ``_FakeRepo`` / ``_frame`` 同时服务「行业豁免」与「行业排除」两组测试,
|
||
# 所以单独放在这里,不属于任何一组。
|
||
# ---------------------------------------------------------------------------
|
||
|
||
|
||
class _FakeRepo:
|
||
#: 行业黑名单自检要用的取值域(与 _frame 里的 industry 列保持同源)
|
||
_DEFAULT_INDUSTRIES = [
|
||
"银行", "煤炭开采", "白酒", "全国地产", "区域地产", "房产服务", "园区开发",
|
||
]
|
||
|
||
def __init__(
|
||
self,
|
||
st: set[str] | None = None,
|
||
susp: set[str] | None = None,
|
||
industries: list[str] | None = None,
|
||
) -> None:
|
||
self._st = st or set()
|
||
self._susp = susp or set()
|
||
self._industries = list(
|
||
self._DEFAULT_INDUSTRIES if industries is None else industries
|
||
)
|
||
|
||
def st_symbols(self, asof, include_delisting=True): # noqa: ANN001, ARG002
|
||
return self._st
|
||
|
||
def suspended_on(self, asof): # noqa: ANN001, ARG002
|
||
return self._susp
|
||
|
||
def stock_master(self) -> pd.DataFrame:
|
||
"""只为 ``MarketFilter`` 的行业名自检提供 ``industry`` 取值域。"""
|
||
return pd.DataFrame({"industry": self._industries})
|
||
|
||
|
||
def _frame(**kwargs) -> pd.DataFrame:
|
||
base = {
|
||
"symbol": ["600036.SH", "601088.SH"],
|
||
"name": ["招商银行", "中国神华"],
|
||
"industry": ["银行", "煤炭开采"],
|
||
"market": ["主板", "主板"],
|
||
"exchange": ["SSE", "SSE"],
|
||
"listed_years": [30.0, 17.0],
|
||
"is_fresh": [True, True],
|
||
"total_mv": [8.6e11, 8.8e11],
|
||
"circ_mv": [7.0e11, 7.3e11],
|
||
"avg_amount": [1e9, 8e8],
|
||
"close": [34.19, 44.37],
|
||
"debt_ratio": [0.902, 0.236],
|
||
"total_hldr_eqy_exc_min_int": [3.5e11, 4.0e11],
|
||
}
|
||
base.update(kwargs)
|
||
return pd.DataFrame(base)
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# 滤网:行业排除清单(黑名单)
|
||
#
|
||
# 回归背景:``backtest.yml: universe_exclusions.industries`` 写在库里,
|
||
# 但数据库**没有**「房地产业」这个取值 —— 它被拆成 全国地产 / 区域地产 /
|
||
# 房产服务 / 园区开发。写错名字不会报任何错,只是「一只都没排除」。
|
||
# ---------------------------------------------------------------------------
|
||
|
||
|
||
def _market_filter(exclude: list[str]):
|
||
from hdiv.core.config import MarketFilterConfig
|
||
from hdiv.universe.filters.market import MarketFilter
|
||
|
||
return MarketFilter(MarketFilterConfig(), exclude_industries=exclude)
|
||
|
||
|
||
def test_market_filter_excludes_listed_industries() -> None:
|
||
f = _market_filter(["全国地产", "区域地产"])
|
||
out = f.compute(
|
||
_frame(industry=["区域地产", "煤炭开采"]), _FakeRepo(), date(2024, 6, 28)
|
||
)
|
||
assert bool(out.passed.iloc[0]) is False, "区域地产必须被排除"
|
||
assert bool(out.passed.iloc[1]) is True, "未列入黑名单的行业不受影响"
|
||
assert "排除清单" in out.reasons["600036.SH"]
|
||
|
||
|
||
def test_industry_exclusion_is_the_reported_reason() -> None:
|
||
"""同时市值不足时,报出的必须是「行业被排除」。
|
||
|
||
若行业判定排在市值之后,被排除的股票会先以「市值不足」落选,
|
||
事后无法分辨「这个行业不做了」还是「这只真的不达标」。
|
||
"""
|
||
f = _market_filter(["全国地产"])
|
||
df = _frame(industry=["全国地产", "煤炭开采"], total_mv=[1e10, 8.8e11])
|
||
out = f.compute(df, _FakeRepo(), date(2024, 6, 28))
|
||
assert bool(out.passed.iloc[0]) is False
|
||
assert "排除清单" in out.reasons["600036.SH"]
|
||
assert "市值" not in out.reasons["600036.SH"]
|
||
|
||
|
||
def test_empty_exclusion_list_changes_nothing() -> None:
|
||
"""不配 = 与改动前逐字一致(本清单只做减法)。"""
|
||
out = _market_filter([]).compute(
|
||
_frame(industry=["全国地产", "煤炭开采"]), _FakeRepo(), date(2024, 6, 28)
|
||
)
|
||
assert out.passed.all()
|
||
|
||
|
||
def test_industry_exclusion_allows_all_four_real_estate_labels() -> None:
|
||
"""四个地产口径都要能被单独命中(配置里缺一个就少排一类)。"""
|
||
labels = ["全国地产", "区域地产", "房产服务", "园区开发"]
|
||
f = _market_filter(labels)
|
||
out = f.compute(
|
||
_frame(industry=["房产服务", "园区开发"]), _FakeRepo(), date(2024, 6, 28)
|
||
)
|
||
assert not out.passed.any()
|
||
|
||
|
||
def test_unknown_industry_name_raises_instead_of_silently_passing() -> None:
|
||
"""写错行业名(库里不存在的「房地产业」)必须报错。
|
||
|
||
这是本功能最容易踩的坑:名单写错时股票池看起来「排除了」,
|
||
实际一只没少 —— 静默失效。宁可直接失败。
|
||
"""
|
||
from hdiv.core.errors import ConfigError
|
||
|
||
f = _market_filter(["房地产业"])
|
||
with pytest.raises(ConfigError) as ei:
|
||
f.compute(_frame(), _FakeRepo(), date(2024, 6, 28))
|
||
msg = str(ei.value)
|
||
assert "房地产业" in msg
|
||
# 必须给出最接近的真实取值,否则用户只能自己去翻库
|
||
assert "全国地产" in msg, f"错误信息没给出候选:{msg}"
|
||
|
||
|
||
def test_risk_filter_exempts_banks_from_leverage() -> None:
|
||
from hdiv.core.config import RiskFilterConfig
|
||
from hdiv.universe.filters.risk import RiskFilter
|
||
|
||
f = RiskFilter(RiskFilterConfig(max_debt_to_assets=0.80), exempt_leverage=["银行"])
|
||
out = f.compute(_frame(), _FakeRepo(), date(2024, 6, 28))
|
||
assert bool(out.passed.iloc[0]) is True, "银行必须豁免负债率上限"
|
||
assert bool(out.passed.iloc[1]) is True, "低负债公司自然通过"
|
||
assert out.values["600036.SH"]["debt_ratio_exempt"] is True
|
||
|
||
|
||
def test_risk_filter_still_rejects_high_leverage_non_exempt() -> None:
|
||
from hdiv.core.config import RiskFilterConfig
|
||
from hdiv.universe.filters.risk import RiskFilter
|
||
|
||
f = RiskFilter(RiskFilterConfig(max_debt_to_assets=0.80), exempt_leverage=["银行"])
|
||
df = _frame(industry=["房地产", "煤炭开采"])
|
||
out = f.compute(df, _FakeRepo(), date(2024, 6, 28))
|
||
assert bool(out.passed.iloc[0]) is False
|
||
assert "资产负债率" in out.reasons["600036.SH"]
|
||
|
||
|
||
def test_quality_filter_exempts_banks_from_fcf_and_leverage() -> None:
|
||
from hdiv.core.config import QualityFilterConfig
|
||
from hdiv.universe.filters.quality import FinancialQualityFilter
|
||
|
||
cfg = QualityFilterConfig(min_ocf_to_profit=0.60, max_debt_to_assets=0.80)
|
||
f = FinancialQualityFilter(cfg, exempt_leverage=["银行"], exempt_fcf=["银行"])
|
||
df = _frame(roe_avg=[0.1513, 0.1406], ocf_to_profit=[-0.03, 1.80])
|
||
out = f.compute(df, _FakeRepo(), date(2024, 6, 28))
|
||
assert bool(out.passed.iloc[0]) is True, "银行豁免 FCF 与负债率"
|
||
|
||
|
||
def test_quality_filter_uses_annual_average_not_quarterly() -> None:
|
||
"""季报 ROE 3.47% 不该被拿去比年均 8% 的阈值。"""
|
||
from hdiv.core.config import QualityFilterConfig
|
||
from hdiv.universe.filters.quality import FinancialQualityFilter
|
||
|
||
cfg = QualityFilterConfig(min_roe_5y_avg=0.08, min_ocf_to_profit=None)
|
||
f = FinancialQualityFilter(cfg)
|
||
df = _frame(
|
||
industry=["白酒", "煤炭开采"],
|
||
roe=[0.0347, 0.0380], # 季报累计值(会被 avg 覆盖)
|
||
roe_avg=[0.1513, 0.1406], # 年报 5 年平均
|
||
)
|
||
out = f.compute(df, _FakeRepo(), date(2024, 6, 28))
|
||
assert bool(out.passed.iloc[0]) is True, "应使用 roe_avg 而非季报 roe"
|
||
assert out.values["600036.SH"]["roe"] == pytest.approx(0.1513)
|
||
|
||
|
||
def test_quality_filter_falls_back_to_latest_when_no_average() -> None:
|
||
from hdiv.core.config import QualityFilterConfig
|
||
from hdiv.universe.filters.quality import FinancialQualityFilter
|
||
|
||
cfg = QualityFilterConfig(min_roe_5y_avg=0.08, min_ocf_to_profit=None)
|
||
f = FinancialQualityFilter(cfg)
|
||
df = _frame(industry=["白酒", "煤炭开采"], roe=[0.20, 0.03]) # 无 roe_avg 列
|
||
out = f.compute(df, _FakeRepo(), date(2024, 6, 28))
|
||
assert bool(out.passed.iloc[0]) is True
|
||
assert bool(out.passed.iloc[1]) is False
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# 分红连续性:一年宽限期
|
||
# ---------------------------------------------------------------------------
|
||
|
||
|
||
def _div(symbol: str, end_year: int, ex_year: int, dps: float = 1.0, month: int = 7) -> dict:
|
||
return {
|
||
"symbol": symbol,
|
||
"end_date": date(end_year, 12, 31),
|
||
"imp_ann_date": date(ex_year, month - 1 if month > 1 else 12, 1),
|
||
"div_proc": "实施",
|
||
"cash_div_tax": dps,
|
||
"cash_div": dps,
|
||
"stk_div": None,
|
||
"base_share": 10000.0,
|
||
"ex_date": date(ex_year, month, 15),
|
||
}
|
||
|
||
|
||
def test_duplicate_dividend_records_count_once() -> None:
|
||
"""同一 (symbol, ex_date) 的重复记录不得把年度 DPS / 总分红重复累加。
|
||
|
||
年度 DPS 决定 ``dps_cagr_5y`` 与 ``dps_volatility``,总现金分红决定
|
||
``payout_ratio`` 与 ``fcf_dividend_cover``(并直接决定选股),
|
||
因此重复记录必须与 ``ttm_dps_series`` 用同一份聚合口径。
|
||
"""
|
||
from hdiv.core.config import DividendFilterConfig
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
cfg = DividendFilterConfig()
|
||
recs = [
|
||
_div("600036.SH", 2023, 2024, dps=1.0, month=7),
|
||
_div("600036.SH", 2023, 2024, dps=1.0, month=7), # 重复公告
|
||
]
|
||
s = DividendFilter._stats(recs, target_year=2023, asof=date(2024, 12, 31), cfg=cfg)
|
||
assert s["dps_by_year"][2023] == pytest.approx(1.0), "年度 DPS 被重复累加"
|
||
|
||
row = pd.Series({"n_income_attr_p": 5.0e9, "free_cashflow": 8.0e9})
|
||
p = DividendFilter._payout_and_cover(recs, row, date(2024, 12, 31), cfg=cfg)
|
||
# 总现金分红 = 每股 1.0 元 × 基准股本 10000 万股 = 1.0e8
|
||
assert p["total_cash_dividend"] == pytest.approx(1.0 * 10000.0 * 1e4)
|
||
assert p["payout_ratio"] == pytest.approx(1.0 * 10000.0 * 1e4 / 5.0e9)
|
||
|
||
|
||
def test_continuity_within_target_year() -> None:
|
||
"""FY2023 分红已在 2024-05 除权 → target=2023,连续 5 年。"""
|
||
from hdiv.core.config import DividendFilterConfig
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
recs = {2023: 2024, 2022: 2023, 2021: 2022, 2020: 2021, 2019: 2020}
|
||
rows = [_div("600036.SH", fy, ex, month=5) for fy, ex in recs.items()]
|
||
s = DividendFilter._stats(rows, target_year=2023, asof=date(2024, 6, 28),
|
||
cfg=DividendFilterConfig())
|
||
assert s["dividend_continuity_years"] == 5
|
||
assert s["continuity_grace_used"] is False
|
||
|
||
|
||
def test_continuity_grace_when_ex_date_lags() -> None:
|
||
"""神华实测情形:FY2023 分红要 2024-07 才除权,asof=2024-06-28 时不可见。
|
||
|
||
此时必须用一年宽限期从 FY2022 起算,而不是判定为「中断分红」。
|
||
"""
|
||
from hdiv.core.config import DividendFilterConfig
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
rows = [_div("601088.SH", fy, fy + 1, month=7) for fy in (2022, 2021, 2020, 2019, 2018)]
|
||
s = DividendFilter._stats(rows, target_year=2023, asof=date(2024, 6, 28),
|
||
cfg=DividendFilterConfig())
|
||
assert s["latest_dividend_year"] == 2022
|
||
assert s["continuity_start_year"] == 2022
|
||
assert s["continuity_grace_used"] is True
|
||
assert s["dividend_continuity_years"] == 5, "不应因为除权晚而误判中断"
|
||
|
||
|
||
def test_continuity_zero_when_genuinely_stopped() -> None:
|
||
"""最近可见分红比应考核财年早两年以上 → 视为真的中断。"""
|
||
from hdiv.core.config import DividendFilterConfig
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
rows = [_div("000002.SZ", fy, fy + 1) for fy in (2019, 2018, 2017, 2016, 2015)]
|
||
s = DividendFilter._stats(rows, target_year=2023, asof=date(2024, 6, 28),
|
||
cfg=DividendFilterConfig())
|
||
assert s["dividend_continuity_years"] == 0
|
||
|
||
|
||
def test_continuity_breaks_on_gap() -> None:
|
||
"""有断档:2022 有、2021 无 → 连续 1 年。"""
|
||
from hdiv.core.config import DividendFilterConfig
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
rows = [_div("X.SZ", fy, fy + 1) for fy in (2022, 2020, 2019)]
|
||
s = DividendFilter._stats(rows, target_year=2023, asof=date(2024, 6, 28),
|
||
cfg=DividendFilterConfig())
|
||
assert s["dividend_continuity_years"] == 1
|
||
|
||
|
||
def test_ttm_dps_only_counts_ex_date_in_window() -> None:
|
||
from hdiv.core.config import DividendFilterConfig
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
rows = [
|
||
_div("X.SZ", 2023, 2024, dps=1.5, month=4), # 窗口内
|
||
_div("X.SZ", 2022, 2023, dps=1.2, month=7), # 窗口内(>2023-06-28)
|
||
_div("X.SZ", 2021, 2022, dps=1.0, month=7), # 窗口外
|
||
]
|
||
s = DividendFilter._stats(rows, target_year=2023, asof=date(2024, 6, 28),
|
||
cfg=DividendFilterConfig())
|
||
assert s["ttm_dps"] == pytest.approx(2.7)
|
||
|
||
|
||
def test_shift_year_handles_leap_day() -> None:
|
||
from hdiv.universe.filters.dividend import _shift_year
|
||
|
||
assert _shift_year(date(2024, 2, 29), -1) == date(2023, 2, 28)
|
||
assert _shift_year(date(2024, 6, 28), -1) == date(2023, 6, 28)
|
||
|
||
|
||
def test_target_year_uses_latest_annual_report() -> None:
|
||
"""最新年报为 FY2023 时,考核目标年应为 2023。"""
|
||
from hdiv.core.config import DividendFilterConfig
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
df = pd.DataFrame(
|
||
{
|
||
"symbol": ["A.SH", "B.SH"],
|
||
"fin_end_date": [date(2024, 3, 31), date(2023, 12, 31)],
|
||
}
|
||
)
|
||
t = DividendFilter._target_years(df, date(2024, 6, 28))
|
||
# 能看到 2024Q1 报,说明 FY2023 年报必然已披露 → 目标年 2023
|
||
assert t["A.SH"] == 2023, "有 2024Q1 报 → FY2023 年报已出"
|
||
assert t["B.SH"] == 2023, "有 FY2023 年报 → 目标是 2023"
|
||
|
||
df2 = pd.DataFrame({"symbol": ["C.SH"], "fin_end_date": [date(2023, 9, 30)]})
|
||
assert DividendFilter._target_years(df2, date(2024, 6, 28))["C.SH"] == 2022, (
|
||
"只看到 2023Q3 → FY2023 年报未出,退回到 2022"
|
||
)
|
||
|
||
|
||
def test_payout_ratio_uses_base_share() -> None:
|
||
"""总现金分红 = 每股分红 × 基准股本(万股→股)。"""
|
||
from hdiv.core.config import DividendFilterConfig
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
rows = [
|
||
{
|
||
"symbol": "X.SH", "end_date": date(2023, 12, 31), "ex_date": date(2024, 6, 1),
|
||
"div_proc": "实施", "cash_div_tax": 2.0, "base_share": 10000.0, # 1 亿股
|
||
}
|
||
]
|
||
row = pd.Series({"n_income_attr_p": 4e8, "free_cashflow": 8e8})
|
||
out = DividendFilter._payout_and_cover(rows, row, date(2024, 6, 28),
|
||
DividendFilterConfig())
|
||
assert out["total_cash_dividend"] == pytest.approx(2.0 * 10000.0 * 1e4)
|
||
assert out["payout_ratio"] == pytest.approx(0.5)
|
||
# 总现金分红 = 2.0 元 × 10000 万股 × 1e4 = 2e8 元;FCF 8e8 → 覆盖 4 倍
|
||
assert out["fcf_dividend_cover"] == pytest.approx(4.0)
|
||
|
||
|
||
def test_fcf_fallback_when_tushare_missing() -> None:
|
||
from hdiv.core.config import DividendFilterConfig
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
rows = [{
|
||
"symbol": "X.SH", "end_date": date(2023, 12, 31), "ex_date": date(2024, 6, 1),
|
||
"div_proc": "实施", "cash_div_tax": 1.0, "base_share": 1000.0,
|
||
}]
|
||
row = pd.Series({
|
||
"n_income_attr_p": 5e6, "free_cashflow": None,
|
||
"n_cashflow_act": 1e7, "c_pay_dist_dpcp_int_exp": 2e6,
|
||
})
|
||
out = DividendFilter._payout_and_cover(rows, row, date(2024, 6, 28),
|
||
DividendFilterConfig())
|
||
assert out["fcf_source"] == "ocf_minus_dist"
|
||
assert out["free_cashflow"] == pytest.approx(8e6)
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# 滤网结果索引契约(曾经把 DataFrame 索引当成 symbol 用的真实 bug)
|
||
# ---------------------------------------------------------------------------
|
||
|
||
|
||
def test_filter_outcome_index_is_dataframe_index_not_symbol() -> None:
|
||
"""滤网的 passed 索引必须与传入 DataFrame 的索引一致。
|
||
|
||
选择器据此用 ``live.at[i, "symbol"]`` 映射;
|
||
若误把索引当 symbol,会导致「全部淘汰」(本项目开发中确实发生过)。
|
||
"""
|
||
from hdiv.core.config import MarketFilterConfig
|
||
from hdiv.universe.filters.market import MarketFilter
|
||
|
||
df = _frame()
|
||
df.index = [10, 20] # 非默认索引
|
||
f = MarketFilter(MarketFilterConfig(min_market_cap=1e10))
|
||
out = f.compute(df, _FakeRepo(), date(2024, 6, 28))
|
||
assert list(out.passed.index) == [10, 20]
|
||
assert set(out.passed.values) <= {True, False}
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# 数据缺失策略(安全边际策略的关键取舍)
|
||
# ---------------------------------------------------------------------------
|
||
|
||
|
||
def _div_cfg(**kw):
|
||
from hdiv.core.config import DividendFilterConfig
|
||
|
||
base = {
|
||
"min_dividend_yield": None,
|
||
"min_continuous_years": 0,
|
||
"min_dividend_years_in_window": 0,
|
||
"max_payout_ratio": None,
|
||
"require_positive_fcf": True,
|
||
"min_fcf_dividend_cover": None,
|
||
}
|
||
base.update(kw)
|
||
return DividendFilterConfig(**base)
|
||
|
||
|
||
def test_missing_fcf_passes_by_default() -> None:
|
||
"""默认宽松:数据缺失放行,避免因未同步而误杀。"""
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
cfg = _div_cfg(on_missing_data="pass")
|
||
stats = {"free_cashflow": None, "dividend_continuity_years": 0,
|
||
"dividend_years_in_window": 0, "dividend_yield": 0.05}
|
||
assert DividendFilter._reject_reason(stats, cfg) is None
|
||
|
||
|
||
def test_missing_fcf_rejected_when_strict() -> None:
|
||
"""严格模式:无法验证现金流即淘汰 —— 忠于「安全边际」的策略逻辑。"""
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
cfg = _div_cfg(on_missing_data="reject")
|
||
stats = {"free_cashflow": None, "dividend_continuity_years": 0,
|
||
"dividend_years_in_window": 0, "dividend_yield": 0.05}
|
||
why = DividendFilter._reject_reason(stats, cfg)
|
||
assert why is not None and "缺失" in why
|
||
|
||
|
||
def test_negative_fcf_rejected_in_both_modes() -> None:
|
||
"""FCF 为负时两种模式都必须淘汰 —— 宽松不等于放行已知风险。"""
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
stats = {"free_cashflow": -1e8, "dividend_continuity_years": 0,
|
||
"dividend_years_in_window": 0, "dividend_yield": 0.05}
|
||
for mode in ("pass", "reject"):
|
||
why = DividendFilter._reject_reason(stats, _div_cfg(on_missing_data=mode))
|
||
assert why is not None and "负" in why, f"{mode} 模式必须拒绝负 FCF"
|
||
|
||
|
||
def test_missing_coverage_strict_only() -> None:
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
stats = {"free_cashflow": 1e8, "fcf_dividend_cover": None,
|
||
"dividend_continuity_years": 0, "dividend_years_in_window": 0,
|
||
"dividend_yield": 0.05}
|
||
assert DividendFilter._reject_reason(
|
||
stats, _div_cfg(min_fcf_dividend_cover=1.0, on_missing_data="pass")
|
||
) is None
|
||
assert DividendFilter._reject_reason(
|
||
stats, _div_cfg(min_fcf_dividend_cover=1.0, on_missing_data="reject")
|
||
) is not None
|
||
|
||
|
||
def test_on_missing_data_invalid_value_rejected() -> None:
|
||
from hdiv.core.errors import SchemaValidationError
|
||
|
||
with pytest.raises(Exception) as ei:
|
||
_div_cfg(on_missing_data="maybe")
|
||
assert "on_missing_data" in str(ei.value) or "Input should be" in str(ei.value)
|
||
del SchemaValidationError
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# 分红支付率必须与分红**同财年**(真实踩到的错误)
|
||
# ---------------------------------------------------------------------------
|
||
|
||
|
||
def test_payout_uses_same_fiscal_year_not_latest_quarter() -> None:
|
||
"""回归:曾用「FY2023 分红 ÷ 2024Q1 净利润」算出 230.9% 的荒谬支付率。
|
||
|
||
正确口径下美的集团 FY2023 为:分红 207.8 亿 ÷ 净利 337.2 亿 = 61.63%。
|
||
"""
|
||
from hdiv.core.config import DividendFilterConfig
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
recs = [{
|
||
"symbol": "000333.SZ", "end_date": date(2023, 12, 31),
|
||
"ex_date": date(2024, 5, 15), "div_proc": "实施",
|
||
"cash_div_tax": 3.0, "base_share": 692675.9241,
|
||
}]
|
||
# 同财年(FY2023)财务
|
||
fy_row = pd.Series({
|
||
"n_income_attr_p": 3.372e10, "free_cashflow": 7.16e10,
|
||
})
|
||
out = DividendFilter._payout_and_cover(
|
||
recs, fy_row, date(2024, 6, 28), DividendFilterConfig()
|
||
)
|
||
assert out["financial_year"] == 2023
|
||
assert out["payout_ratio"] == pytest.approx(207.8 / 337.2, abs=0.01)
|
||
assert out["payout_basis"] == "same_fiscal_year"
|
||
|
||
|
||
def test_payout_not_computed_without_same_year_row() -> None:
|
||
"""缺少同财年财务时必须返回「不可得」,而不是退回到最新季报。"""
|
||
from hdiv.core.config import DividendFilterConfig
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
recs = [{
|
||
"symbol": "X.SZ", "end_date": date(2023, 12, 31),
|
||
"ex_date": date(2024, 5, 15), "div_proc": "实施",
|
||
"cash_div_tax": 1.0, "base_share": 10000.0,
|
||
}]
|
||
out = DividendFilter._payout_and_cover(
|
||
recs, None, date(2024, 6, 28), DividendFilterConfig()
|
||
)
|
||
assert out["payout_ratio"] is None
|
||
assert out["fcf_dividend_cover"] is None
|
||
assert out["payout_basis"] == "missing_same_year_financials"
|
||
|
||
|
||
def test_total_cash_dividend_uses_base_share() -> None:
|
||
"""base_share 是**万股**,漏乘 1e4 会把支付率缩小一万倍。"""
|
||
from hdiv.core.config import DividendFilterConfig
|
||
from hdiv.universe.filters.dividend import DividendFilter
|
||
|
||
recs = [{
|
||
"symbol": "X.SZ", "end_date": date(2023, 12, 31),
|
||
"ex_date": date(2024, 5, 15), "div_proc": "实施",
|
||
"cash_div_tax": 2.0, "base_share": 10000.0, # 1 亿股
|
||
}]
|
||
row = pd.Series({"n_income_attr_p": 4e8, "free_cashflow": 8e8})
|
||
out = DividendFilter._payout_and_cover(
|
||
recs, row, date(2024, 6, 28), DividendFilterConfig()
|
||
)
|
||
assert out["total_cash_dividend"] == pytest.approx(2.0 * 10000.0 * 1e4)
|
||
assert out["payout_ratio"] == pytest.approx(0.5)
|
||
|
||
|
||
def test_dividend_records_include_base_share() -> None:
|
||
"""回归:repo.dividend_records 曾漏选 base_share,
|
||
|
||
导致 payout_ratio 与 fcf_dividend_cover 在全库范围内静默为 NULL,
|
||
进而使 max_payout_ratio / min_fcf_dividend_cover 两个筛选条件从未生效。
|
||
"""
|
||
import inspect
|
||
|
||
from hdiv.data.repo import Repo
|
||
|
||
src = inspect.getsource(Repo.dividend_records)
|
||
assert "base_share" in src, "dividend_records 必须选出 base_share"
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# 重跑覆盖同一条记录
|
||
# ---------------------------------------------------------------------------
|
||
|
||
|
||
def test_universe_run_id_is_deterministic() -> None:
|
||
"""回归:run_id 不得含时间戳,否则同参数重跑会不断累积重复记录。
|
||
|
||
早期实现把 datetime.now() 编进指纹,同一 asof 最多累积了 11 条内容相同的记录。
|
||
现在的语义是「同一份配置 + 同一时点 → 同一个 run_id → 重跑原地覆盖」。
|
||
"""
|
||
import inspect
|
||
|
||
from hdiv.universe import selector
|
||
|
||
src = inspect.getsource(selector.UniverseSelector)
|
||
i = src.find("run_id = stable_id(")
|
||
assert i != -1, "未找到 run_id 生成处"
|
||
# 取到该语句结束的分号行(不能用第一个 ')',那会截断在 config_hash(self.config) 里)
|
||
end = src.find("\n )", i)
|
||
assert end != -1, "未找到 run_id 语句结尾"
|
||
block = src[i:end]
|
||
assert "datetime.now" not in block, f"run_id 指纹仍含时间戳:{block}"
|
||
for must in ("config_hash", "effective", "self.config.name"):
|
||
assert must in block, f"run_id 指纹缺少 {must}:{block}"
|
||
|
||
|
||
@pytest.mark.db
|
||
def test_universe_rerun_overwrites_same_record() -> None:
|
||
"""同参数重跑不新增记录,且成员行数等于候选数(无重复堆积)。"""
|
||
from hdiv.core.config import load_config
|
||
from hdiv.data import db
|
||
from hdiv.data.sync.base import stable_id
|
||
|
||
db.load_dotenv_once()
|
||
cfg = load_config("datasource")
|
||
df = db.read_sql(
|
||
"SELECT r.run_id, r.candidate_count, r.asof_date, r.config_hash, r.name, "
|
||
" COUNT(m.id) AS member_rows "
|
||
"FROM hd_universe_run r LEFT JOIN hd_universe_member m ON m.run_id = r.run_id "
|
||
"GROUP BY r.run_id HAVING member_rows > 0 "
|
||
"ORDER BY r.created_at DESC LIMIT 5",
|
||
cfg=cfg,
|
||
)
|
||
if df.empty:
|
||
pytest.skip("没有筛选记录")
|
||
checked = 0
|
||
for _, r in df.iterrows():
|
||
expect = stable_id("universe", r["name"], str(r["asof_date"]), r["config_hash"])
|
||
if r["run_id"] != expect:
|
||
continue # 确定化之前的历史记录,跳过
|
||
checked += 1
|
||
assert int(r["member_rows"]) == int(r["candidate_count"]), (
|
||
f"run_id={r['run_id'][:10]} 成员行数 {r['member_rows']} "
|
||
f"应等于候选数 {r['candidate_count']}(出现重复堆积)"
|
||
)
|
||
assert checked > 0, "未找到确定化之后生成的筛选记录,无法验证"
|