feat: 增量处理与 pipeline 断点续跑
- M2 run_extractor: 产物存在即跳过提取(仅回补 index),--force 全量;
增量成功率统计含跳过项,修复全跳过时误报失败
- M4 run_event_extraction: data/events/{day}/{url_hash}.json 已存在即跳过,
不重复调用 LLM API;--force 全量;failed 只留本次失败
- M5 run_embedding: data/embeddings/{day}/{url_hash}.json 已存在即跳过,
不重复调用 embed API;--force 全量
- scheduler/pipeline: 步骤结果按日期写入 data/pipeline/state.json(原子写),
run_pipeline(resume=True) 从首个失败/未执行步骤续跑
- run_scheduler + a-share CLI: --once --resume 断点续跑(--steps 互斥)
- 新增 tests/test_incremental.py 10 个测试(跳过逻辑 + resume)
- .gitignore: 忽略 data/pipeline/ 运行状态
This commit is contained in:
@@ -59,11 +59,17 @@ def _parse_schedule_times(raw: str) -> list[tuple[int, int]]:
|
||||
|
||||
|
||||
def _once(args: argparse.Namespace) -> int:
|
||||
"""单次执行模式。"""
|
||||
"""单次执行模式。
|
||||
|
||||
默认全量执行;--resume 时断点续跑(跳过连续成功步骤,从失败/未执行步骤继续)。
|
||||
"""
|
||||
steps = None
|
||||
if args.steps:
|
||||
steps = [s.strip() for s in args.steps.split(",")]
|
||||
run_pipeline(args.date, steps=steps)
|
||||
if args.resume and args.steps:
|
||||
logger.error("--resume 与 --steps 不能同时使用(断点续跑针对全链路)")
|
||||
return 2
|
||||
run_pipeline(args.date, steps=steps, resume=args.resume)
|
||||
return 0
|
||||
|
||||
|
||||
@@ -185,6 +191,10 @@ def main() -> int:
|
||||
)
|
||||
parser.add_argument("--steps", default=None,
|
||||
help="仅执行指定步骤,逗号分隔 (如 crawler,extractor)")
|
||||
parser.add_argument(
|
||||
"--resume", action="store_true",
|
||||
help="断点续跑(仅 --once):跳过连续成功步骤,从上次失败/未执行步骤继续",
|
||||
)
|
||||
parser.add_argument("--log-level", default="INFO")
|
||||
args = parser.parse_args()
|
||||
|
||||
|
||||
Reference in New Issue
Block a user