feat: 增量处理与 pipeline 断点续跑

- M2 run_extractor: 产物存在即跳过提取(仅回补 index),--force 全量;
  增量成功率统计含跳过项,修复全跳过时误报失败
- M4 run_event_extraction: data/events/{day}/{url_hash}.json 已存在即跳过,
  不重复调用 LLM API;--force 全量;failed 只留本次失败
- M5 run_embedding: data/embeddings/{day}/{url_hash}.json 已存在即跳过,
  不重复调用 embed API;--force 全量
- scheduler/pipeline: 步骤结果按日期写入 data/pipeline/state.json(原子写),
  run_pipeline(resume=True) 从首个失败/未执行步骤续跑
- run_scheduler + a-share CLI: --once --resume 断点续跑(--steps 互斥)
- 新增 tests/test_incremental.py 10 个测试(跳过逻辑 + resume)
- .gitignore: 忽略 data/pipeline/ 运行状态
This commit is contained in:
2026-08-12 10:19:27 +08:00
parent c1a803968a
commit 2b4efea219
8 changed files with 518 additions and 35 deletions
+4
View File
@@ -179,6 +179,8 @@ def cmd_pipeline(args: argparse.Namespace) -> int:
return _run_module("scripts.run_scheduler", extra, timeout=pipeline_timeout)
if args.once:
extra = ["--once", "--date", args.date or _today()]
if args.resume:
extra.append("--resume")
steps = args.steps
if args.report:
steps = (steps + ",report") if steps else "report"
@@ -878,6 +880,8 @@ def main() -> int:
p.add_argument("--once", action="store_true", help="立即执行一次")
p.add_argument("--date", default=None)
p.add_argument("--steps", default=None, help="指定步骤(crawler,extractor,...,report)")
p.add_argument("--resume", action="store_true",
help="断点续跑(仅 --once):跳过连续成功步骤,从上次失败/未执行步骤继续")
p.add_argument("--report", action="store_true", help="全链路末尾生成日报")
p.add_argument("--cninfo-once", action="store_true", help="cninfo watchlist 全链路(公告+调研+IRM)")
p.set_defaults(func=cmd_pipeline)