perf(backend): 内存优化三项——全市场研究不再占满 8G

1) 数据装配流式+列裁剪:Repository 新增 stream_range_many_columns(只 SELECT
   所需列、SQL 侧转 REAL、yield_per 分批),引擎按 required_columns 取数
   (LocalEngine 仅 close+因子字段),消除 ORM/Decimal 全量物化;
2) 研究 Job 独立子进程执行(job.mode=subprocess):python -m app.cli.run_job
   在子进程内设 RLIMIT_AS 上限,OOM 归档 failed 而非拖垮 API worker;
   子进程异常退出由父进程补记 failed;并发上限 2;
3) 服务启动清理:残留 queued/running Job 标记 failed(防永久 running)。

实测同款全市场回测:uvicorn worker RSS 稳定 ~220MB,任务峰值内存由 4.1GB+
降至 ~470MB,24s 完成并归档(此前 43s 未完成即 OOM)。
新增/更新测试 96 passed,ruff 干净。
This commit is contained in:
Simon
2026-09-06 22:12:44 +08:00
parent 02e42184be
commit 195f5d41f4
18 changed files with 593 additions and 31 deletions
+9 -2
View File
@@ -36,8 +36,15 @@ storage:
qlib_dir: "data/qlib"
job:
# 第一阶段异步任务模式:local(FastAPI BackgroundTasks 级);Phase 复杂后再引入队列
mode: "local"
# 研究任务执行模式:
# subprocess —— 独立子进程执行(内存隔离 + RLIMIT 上限,防研究任务 OOM 拖垮 API 服务)
# local —— 本进程内执行(开发 / 单测,无隔离)
# 可用环境变量 JOB_MODE 覆盖(测试强制 local)。
mode: "subprocess"
# subprocess 模式下子进程虚拟地址空间上限(GB);达到上限任务以 MemoryError 失败,不会 OOM 整机
max_memory_gb: 6
# 并发研究子进程上限(超出直接失败并提示稍后再试)
max_concurrent_jobs: 2
agent:
llm: