Compare commits
20 Commits
d4617812c5
...
74d09ccbda
| Author | SHA1 | Date | |
|---|---|---|---|
| 74d09ccbda | |||
| 552382bf2d | |||
| a0389e4f33 | |||
| 3082db5a31 | |||
| c4ae457563 | |||
| d51ee7df72 | |||
| 066a2d3b87 | |||
| cf451466a6 | |||
| 275c66baf5 | |||
| 5231dbc99a | |||
| 73bddca583 | |||
| bfd08561fc | |||
| 28e648b4fe | |||
| eaac48a1db | |||
| ca264defbd | |||
| 031348c9a3 | |||
| f45e14234e | |||
| 91c5e2a4a3 | |||
| 73ba705bca | |||
| 86bfbcf283 |
@@ -0,0 +1,52 @@
|
||||
# 夜间进展审计报告(2026-10-06 23:10 定时班)
|
||||
|
||||
- **审计时点**:2026-10-06 23:10(周日休市,国庆连休第 6 天;复市倒数第 2 天)
|
||||
- **基线锚点**:`f1c5ac8`(10-05 22:46 双报告归档)→ **HEAD `327198e`**
|
||||
- **方法**:f1c5ac8..HEAD 仅 1 笔纯文档提交、无 P0/P1 候选,故本轮未派 Explore 子代理,全部主线程直读实证(git show 逐节 + grep 坐标核对);只读铁律遵守,报告为本轮唯一写入。
|
||||
|
||||
---
|
||||
|
||||
## 一、新提交审阅(1 笔,零代码)
|
||||
|
||||
| 提交 | 内容 | 审阅结论 |
|
||||
|---|---|---|
|
||||
| `327198e`(10-06 08:30)[nas] | 代码线终审报告 §八 时差勘误+夜间收编补验落库(audit/20261005_final_verdict.md +54 行,唯一文件) | **通过,无发现**。①内容完整性:监控线终裁 session 的两处勘正与交叉核验尾注**完整入库**(`2^64≈1.8e19` 勘正、`:105/:160` 行号勘正、`二审交叉核验` 尾注各 grep 命中)——用户指差的时差问题就此在档闭环;②标签纪律:`[nas]` 符合 audit/docs 提交惯例(74bced6/f1c5ac8 同款);③纯文档零码面,无 CI 门禁影响、无敏感信息;④commit message 与实际内容逐项相符(含「MON-202610-10 仍欠」「本机 4 测试失败系缺 polars 非回归」两条诚实注记)。 |
|
||||
|
||||
**工作区**:`git status` 全清,无未跟踪/未提交件。
|
||||
|
||||
## 二、在途清单进展对照(本日关闭数:0)
|
||||
|
||||
| 项 | 排期 | 状态 | 证据 |
|
||||
|---|---|---|---|
|
||||
| P2-19 探活撤单败灯色 | 10-08 验证窗后 | 仍开放(符合排期) | `scripts/qmt_relogin/` 自 f1c5ac8 零变更(git log 路径过滤空) |
|
||||
| P3-14 双机同日竞态 | 低危在途 | 仍开放 | 同上(data_gap_check 零变更) |
|
||||
| P3-23 rc 分档打标 | 在途 | 仍开放 | 同上 |
|
||||
| P3-25 argv key 窗口(=MON-08) | 10-08 后 --env-file 收敛 | 仍开放 | 同上 |
|
||||
| P3-27 margin 缺块记录项 | 记录维持 | 仍开放 | 同上 |
|
||||
| P3-28/29/30/31 qmt 灾备源族 | 10-08 后随 identity 班车 | 仍开放 | `scripts/qmt_relogin/` 零变更 |
|
||||
| **MON-202610-10 台账补录**(RB-C3 承诺) | 无明确 owner/时点 | **仍欠(第三次在案)** | monitoring-design §11 grep `MON-202610-10`/`integrity_gate` 零命中(今晨终审 §八尾注点名后仍未动)|
|
||||
| MON-06 always 族 / MON-09 morn 同峰 | 10-08 验证窗后 | 未被提前处置 ✓(提前动调度=异常,未发生) | 调度/注册文件零变更 |
|
||||
| ⚖️-3 monitor_common.py 收口 | 10-08 后三域合议 | 未启动 ✓(符合排期) | `sanguo_data/monitor_common.py` 不存在 |
|
||||
| funnel LLM 开闸 | 待令 | **闸仍未开** ✓ | `run_corpus_standalone.sh:116` 仍带 `--no-llm`(:102 B 案注释在位) |
|
||||
|
||||
**结论:零代码提交、零失管、零提前处置;唯一持续欠账=MON-202610-10(终审 §五.3 与 §八尾注两次点名,尚无认领记录)。**
|
||||
|
||||
## 三、10-08 复市验证窗前瞻(文档面就绪度:4/4 就绪)
|
||||
|
||||
| 验证项 | 文档锚点 | 就绪 |
|
||||
|---|---|---|
|
||||
| probe0915 假期 rc=1 → 复市日 09:15 转绿=桥复市金标准 | spec §3.4:114 + §10 值班表 | ✓ |
|
||||
| morn logs_stale 复市日一班诚实黄(勿处置) | spec §4.2-5⑤:161 + §9.A-15:417(**T-3 契约钉测已在库**,79c6b8b) | ✓ |
|
||||
| error_storm 假日族复市自愈(勿手工消警,仍黄才升级) | spec §4.2-4⑤:158 + §9.A-5:397 + §10 | ✓ |
|
||||
| fills_vs_paper v1.2 首次实评(10-08 eve,预期 n/a 绿;零配对 unpaired 可见性已备) | spec §4.2-9⑤:173 | ✓ |
|
||||
|
||||
附带确认:§10 值班手册各形态行、§1 时间线(10-05 morn 首验绿,待 10-06/07 连续复验)、runbook :122/:346 的 P2-18 生产 wrapper 收尾(10-08 窗后改前报备)均在位。
|
||||
|
||||
## 四、发现汇总
|
||||
|
||||
- **P0/P1/P2:无。**
|
||||
- P3 提示(仅一条,重复在案):MON-202610-10 台账补录自 RB-C3 rebuttal(10-05 12:40)起已两次被终审点名仍未落,且无明确认领 owner——建议明确归属(终审 §五.3 建议 infra)后在 §11 落一条即闭。
|
||||
|
||||
---
|
||||
|
||||
**处置声明:本轮处置均未执行,git 状态保持原状**(HEAD=327198e 工作区 clean;本报告为唯一新增写入)。
|
||||
@@ -0,0 +1,47 @@
|
||||
# 终审验收报告(2026-10-07 晚间班)
|
||||
|
||||
- **执行说明**:原定 23:10 定时任务因平台规则未能创建(本会话即定时任务会话,不允许内部再建;10-05 终审同款先例)——改即时执行。最后一笔提交 `508ead3` 落于当日 21:10,验收输入与 23:10 一致;若夜间再有提交,新对话可补排。
|
||||
- **基线**:`327198e`(10-06 定时班锚点)→ **HEAD `508ead3`**,新提交 **14 笔**(10-07 开工波:data/NAS/account 监控长尾三批 + P3-14 + factor #91 批六件 + identity ASCII 生产实案修复 + docs 三笔)。
|
||||
- **方法**:三路 Explore 并行验收 + 主线程对全部 P2 与运营级发现逐条实证复核(`account_monitor.py:44-62/:283-298`、`data_monitor.py:57-64`、runbook:115/promote 链、`corpus_funnel.py:406-414` 均亲读确认);本地复跑 monitoring 八件套。
|
||||
|
||||
---
|
||||
|
||||
## 一、验收总判(声称修复 30 项)
|
||||
|
||||
| 判定 | 计数 | 明细 |
|
||||
|---|---|---|
|
||||
| **已修-通过** | **27** | data 六件全过(含 `_quantile` 与 numpy 线性法 20000 组随机对拍 0 差的独立补证);NAS 五件全过(N-4 尾读滑窗/N-5 17:45 边角/N-6 main 兜底含崩溃键恒入 green_keys/N-7 TOFU+env 覆盖/out 14d 保洁);account 群 A-2/A-3/A-4/A-6/A-7 五件 + 两件冻结落档(§9.A-19 在档,坐标微漂 3 行 cosmetic);P3-14 三子项全过(vps 侧拒开在一切触网之前+防触网 AssertionError 钉);factor #91①-⑥ 六件全过;identity ASCII 化(round-trip 等值经独立实测:`\u` 转义件 `json.load` 还原 == 中文原路径);文档两笔快检过 |
|
||||
| **部分修复** | **2** | ①**P3-23 verdict 打标**:六 lane 接线成立、daily/backfill/gate 映射与源码 rc 语义一致,但 funnel `rc=1` 存在歧义(`corpus_funnel.py:413` else `ctx.rc()` 在"非 stopped+有噪声 failed"时也返 1,会被打成 `fail_loud` 过度分诊)且 flash/xcheck 未细分;②**A-5 查询超时护栏**:查询面已封(fail-closed+测试真钉),但见下 P2 两条 |
|
||||
| **未修** | **0** | — |
|
||||
| **虚报** | **0** | 全部声称与实现相符,无虚报 |
|
||||
|
||||
**排期纪律核查:违规 0。** P2-19、P3-25、P3-27、P3-28/29/30/31(qmt 灾备源族)、MON-06(`"always": None` 原样 data_monitor.py:363)、MON-09、⚖️-3 monitor_common(文件不存在)、funnel `--no-llm`(:153 闸未开)全部未动,与"10-08 验证窗后"排期一致;e8902c2 提交信息自陈"标尺=应急班车不污染 10-08 归因",两件冻结(traded_at LIKE/wmic 时区)即该纪律的落档。
|
||||
|
||||
## 二、终审新增发现
|
||||
|
||||
**P1(运营提醒,非代码缺陷)——identity 生效位部署链硬期限**:
|
||||
`5d5a9f2` 修的是 repo 侧 `config/qmt_identity.json`,而生效位 **随 promote 走**(runbook:115 + promote.sh ALL_MODS 含 config,主线程复核属实)。identity env_plan §五②把 VPS 班车排"10-08 复市验证窗口后",与 §五①′ 自述"部署须在 **10-08 21:00(xt-daily)前**落地"存在时序张力——**不提前发一次 promote(`--module config` 亦可),明晚 xt_eod 将精确复刻 10-06 事故**(PS5.1 GBK 读非 ASCII json → FormatException → account_id 缺失 → universe=0 → rc=1;30d 窗口使数据可自愈,但当晚班红+当日 ETF/基金增量缺)。**建议:10-08 21:00 前完成一次 promote 发车,并在 cron.log 核对 funnel/xt 段正常。**
|
||||
|
||||
**P2 ×2(A-5 残余,主线程已复核)**:
|
||||
1. `_call_with_timeout` 超时即弃线程(`account_monitor.py:44-62`):每查询新建 daemon 线程、超时弃等——QMT 持续挂死场景 ≈1440 线程/天累积无上界;且 `:46` docstring "守护线程泄漏有界"与实现不符(该句需勘正)。
|
||||
2. `_ensure_trader` 的 `XtQuantTrader()/start()/connect()` 不在护栏内(`account_monitor.py:285-288`)——提交信息自举的"bs.login 挂死同族"恰在此面,同族挂死仍可令 poll 线程永阻。挂死类只封了查询面。
|
||||
处置建议(不执行):连接三连包同一护栏 + docstring 勘正 + 泄漏评估(查询线程数上限/复用 worker)。
|
||||
|
||||
**P3 汇总 ×16(本轮新增,全量坐标见各路验收记录,择要)**:data 侧摄取前崩溃丢事件窗(`os.replace` 与 upsert 间,毫秒级窗口)+崩溃残留 pid tmp 碎件;strategy 域同款 `--db` 字面量脆点仍在(strategy_monitor.py:90,101,超出本批 scope);P3-23 funnel rc=1 歧义+flash/xcheck 未细分+无自动消费方(人工 grep 语义位,按声称成立);NAS TOFU 依赖 OpenSSH≥7.6 未实证(DSM7=8.2p1 推定满足,**部署生效日首跑看 cron.log 无 "Bad SSH option"**);NAS 崩溃件同分钟撞 .pushed marker 极边角;NAS 其余 runner `=no` 裸奔未收口(scope 外);A-4 pid 重用概率撞、stop-during-query C 层行为未证;P3-14 vps 拒开返回 None 与其他跳过不可辨、NAS 权威侧缺班无专属告警(已声明取舍);origin 枚举仅建条路径校验+yaml 手改 typo 静默退闸;np.float32 NaN 纸面暴露;`_append_writeback_missed` 自身 OSError 尾;events 缺省落点 API/CLI 不一致(既有);data_gaps 非 dict 理论面;**第 4 份 Gitea 字面量残留**(strategy_registry.py:282 env 缺省);④钉不守"VPS 生效位手改"路径+round-trip 等值无测试钉(本次人工实证)。
|
||||
|
||||
**遗留未修(维持开放)**:终裁报告 P3-1(data 隔离降级键 `data-{name}-stale` 对 disk/static_exists/vintage 三 kind 绿班永不 resolve,`data_monitor.py:64` 主线程复核原样——NAS 侧同族已防,data 侧未跟);cosmetic:§9.A-19 引 strategy_checks.py:155 实为 :147-154、MON-01 条目内旧"黄 14d"措辞。
|
||||
|
||||
## 三、测试复跑(只读 pytest)
|
||||
|
||||
- monitoring 八件套:**219 passed / 0 failed**(2.13s;较上轮 +19,含修复波新增测试全绿)。
|
||||
- factor/decompose 类本机不可运行(缺 polars,环境缺口非代码问题,与 10-05 终审局限一致,全量以 VPS/CI 为准)。
|
||||
|
||||
## 四、更正后台账
|
||||
|
||||
- **代码线**(终审 §八 基线:关闭 51/挂账 2/开放 11/待办 1):本轮 **P3-14、P3-23 关闭 → 开放 11→9**(P2-19、P3-25、P3-27、P3-28..31 五条排期件 + 记录维持 P3-32/33);挂账 2 不变(⚖️-1 all_weather 窗口)。
|
||||
- **监控线**:MON-10 关闭(05f90f0);文档班车七处全闭(§6.1/§10/§4.0/§5.4②/C-A8/MON-04 标注);P3 长尾 data 6+NAS 5+account 5 关闭;**仍开放**:终裁 P3-1 隔离键、A-5 部分残余 2 条 P2、本轮新增 P3×16、MON-06/09(排期)、NAS 生效位 OpenSSH 实证(部署日动作)、identity promote 发车(运营硬期限)。
|
||||
- **三天审计线总态**:P0=0;代码线无未决 P1;监控线无未决 P1;最高在案代码项=P2×2(A-5 残余);最高在案运营项=P1 提醒(promote 期限,明日 21:00 前)。
|
||||
|
||||
---
|
||||
|
||||
**处置声明:本轮处置均未执行,git 状态保持原状**(HEAD=508ead3 工作区 clean + 本报告与昨夜报告两处 audit/ 新增为仅有的未跟踪件)。
|
||||
@@ -0,0 +1,623 @@
|
||||
# 判定弹药三件+AST 防换皮+test 隔离核实 Implementation Plan(班次一)
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||
|
||||
**Goal:** 给月度批评判定人补三件弹药(全期合成 t/正 IC 占比/逐年 IC)+注册时防换皮提示+quant12_v2a 权重窗隔离审计核实。
|
||||
|
||||
**Architecture:** 聚合源头在 `monthly_review.build_report`(monthly_points 已有逐点 ic_mean/t,只差聚合口径)→ 报告 JSON 新增 `factor_stats` 块 → detail 端点透传 → 前端展示。防换皮=新模块 `expression_match`(Python ast 方言内实现交换律归一化+最大公共子树占比,QuantaAlpha 算法思想不移植解析器)→ decompose 注册路径调用存 `entry.similarity`(提示非硬拒)。审计=读 weight_profiles 档案与考场窗核实无重叠,结论落 spec。
|
||||
|
||||
**Tech Stack:** Python 3.11 stdlib(ast/statistics 零新依赖)、Vue3+Element Plus(既有栈)、pytest/vitest。
|
||||
|
||||
## Global Constraints
|
||||
|
||||
- spec=2026-09-22 流水线总设计 §4.2「判定弹药三件」「AST 防换皮」+§4.4「test 隔离条款」(2026-10-10 补丁块,本次已更新)
|
||||
- **push 纪律**:今晚 21:30 双机日度首班——改动触及 monthly_review(daily 两段链 stage2)。**全部 task 本地 commit 完成,push 等明早首班验证通过后**(用户已知情)。
|
||||
- commit 标签:触及 factors 端点+前端 → `[vps]`(VPS 常驻 api+前端也触发);monthly_review 批链 → `[nas]`。每条 commit 末尾二选一或并存,CI enforce-label 必过。
|
||||
- 测试铁律:凡 decompose/registry 测试注册因子必须 `monkeypatch.setenv("SANGUO_FACTOR_REGISTRY", str(tmp))`;前端 `vue-tsc` 管道退出码显式门控。
|
||||
- 口径注记(plan 内定为规范):monthly point 的 `ic_mean`=该批 12M 窗内日均 IC 均值;跨月点等权均值≈全史等权近似(相邻窗重叠 11/12),月度链根层每月一点(daily/ 子目录不进 monthly_points,判定层月末快照语义不变)。
|
||||
- 阈值常量:相似度提示阈值 `SIM_FLOOR = 0.6`(起步默认,首年校准点)。
|
||||
|
||||
---
|
||||
|
||||
### Task 1: monthly_review 弹药聚合(factor_stats)
|
||||
|
||||
**Files:**
|
||||
- Modify: `sanguo_factor/monthly_review.py`(build_report 内新函数+JSON 字段+render_markdown 增节)
|
||||
- Test: `tests/factor/test_monthly_review.py`(新建,目录已有 conftest.py)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `build_report` 既有 `monthly_points: dict[str, list[dict]]`(每点 `{"month": "YYYY-MM", "t": float|None, "ic_mean": float|None, "count": int}`)
|
||||
- Produces: `factor_stats(points: list[dict]) -> {"icAll": float|None, "tAll": float|None, "positiveRatio": float|None, "byYear": dict[str, float]}`;报告 JSON 顶层新键 `"factor_stats": {name: <上述>}`
|
||||
|
||||
- [ ] **Step 1: 写失败测试**
|
||||
|
||||
```python
|
||||
# tests/factor/test_monthly_review.py
|
||||
"""判定弹药三件聚合口径(spec §4.2 2026-10-10):全期合成t/正IC占比/逐年."""
|
||||
import math
|
||||
|
||||
from sanguo_factor.monthly_review import factor_stats
|
||||
|
||||
|
||||
def _pts(*pairs):
|
||||
return [{"month": m, "ic_mean": v, "t": None, "count": 1} for m, v in pairs]
|
||||
|
||||
|
||||
def test_factor_stats_basic():
|
||||
s = factor_stats(_pts(("2025-09", 0.04), ("2025-10", 0.06), ("2026-09", 0.02)))
|
||||
assert s["icAll"] == round((0.04 + 0.06 + 0.02) / 3, 6)
|
||||
assert s["positiveRatio"] == 1.0
|
||||
assert s["byYear"] == {"2025": round((0.04 + 0.06) / 2, 6), "2026": 0.02}
|
||||
mean = 0.04
|
||||
var = ((0.0) + (0.02) + (-0.02)) / 2 # ddof=1
|
||||
t_all = mean / (math.sqrt(var) / math.sqrt(3))
|
||||
assert s["tAll"] == round(t_all, 4)
|
||||
|
||||
|
||||
def test_factor_stats_negative_and_short():
|
||||
s = factor_stats(_pts(("2025-09", 0.05), ("2025-10", -0.01)))
|
||||
assert s["positiveRatio"] == 0.5
|
||||
assert s["tAll"] is not None
|
||||
single = factor_stats(_pts(("2025-09", 0.05)))
|
||||
assert single["tAll"] is None # n<2 无合成 t
|
||||
assert single["positiveRatio"] == 1.0
|
||||
|
||||
|
||||
def test_factor_stats_empty_and_dirty():
|
||||
assert factor_stats([]) == {"icAll": None, "tAll": None,
|
||||
"positiveRatio": None, "byYear": {}}
|
||||
dirty = [{"month": "2025-09", "ic_mean": None, "t": None},
|
||||
{"month": "", "ic_mean": 0.1, "t": None}]
|
||||
s = factor_stats(dirty)
|
||||
assert s["icAll"] == 0.1
|
||||
assert s["byYear"] == {} # 空 month 不进逐年
|
||||
```
|
||||
|
||||
- [ ] **Step 2: 跑测试确认失败**
|
||||
|
||||
Run: `python3 -m pytest tests/factor/test_monthly_review.py -v`
|
||||
Expected: FAIL `ImportError: cannot import name 'factor_stats'`
|
||||
|
||||
- [ ] **Step 3: 实现 factor_stats + build_report 接线**
|
||||
|
||||
在 `sanguo_factor/monthly_review.py` 的 `build_report` 函数**之前**加入:
|
||||
|
||||
```python
|
||||
def factor_stats(points: list[dict]) -> dict:
|
||||
"""判定弹药聚合(2026-10-10 spec §4.2):全期合成 t/正 IC 占比/逐年.
|
||||
|
||||
口径:各月点 ic_mean(该批 12M 窗日均 IC 均值)的等权均值≈全史近似;
|
||||
合成 t=mean/(std(ddof=1)/√n),n<2 或 std=0 时如实 None(小样本判读弱
|
||||
的诚实注记在 spec,判定动作不因此自动化).
|
||||
"""
|
||||
ics = [p["ic_mean"] for p in points if p.get("ic_mean") is not None]
|
||||
if not ics:
|
||||
return {"icAll": None, "tAll": None, "positiveRatio": None, "byYear": {}}
|
||||
n = len(ics)
|
||||
mean_ic = sum(ics) / n
|
||||
t_all = None
|
||||
if n >= 2:
|
||||
var = sum((v - mean_ic) ** 2 for v in ics) / (n - 1)
|
||||
if var > 0:
|
||||
t_all = round(mean_ic / (var ** 0.5 / n ** 0.5), 4)
|
||||
by_year: dict[str, list[float]] = {}
|
||||
for p in points:
|
||||
m, v = str(p.get("month") or ""), p.get("ic_mean")
|
||||
if m and v is not None:
|
||||
by_year.setdefault(m[:4], []).append(v)
|
||||
return {"icAll": round(mean_ic, 6), "tAll": t_all,
|
||||
"positiveRatio": round(sum(1 for v in ics if v > 0) / n, 4),
|
||||
"byYear": {y: round(sum(v) / len(v), 6)
|
||||
for y, v in sorted(by_year.items())}}
|
||||
```
|
||||
|
||||
`build_report` 返回 dict 处(现含 `"monthly_points": monthly_points` 的字面量)加一行:
|
||||
|
||||
```python
|
||||
"factor_stats": {name: factor_stats(pts)
|
||||
for name, pts in monthly_points.items()},
|
||||
```
|
||||
|
||||
- [ ] **Step 4: 跑测试确认通过**
|
||||
|
||||
Run: `python3 -m pytest tests/factor/test_monthly_review.py -v`
|
||||
Expected: 3 PASS
|
||||
|
||||
- [ ] **Step 5: render_markdown 加弹药节**
|
||||
|
||||
在 `render_markdown` 内(verdicts 表渲染之后、collective 节之前)插入:
|
||||
|
||||
```python
|
||||
stats = report.get("factor_stats") or {}
|
||||
if stats:
|
||||
lines.append("")
|
||||
lines.append("## 判定弹药(全期口径)")
|
||||
lines.append("")
|
||||
lines.append("| 因子 | 全期IC | 合成t | 正IC占比 | 逐年IC |")
|
||||
lines.append("|------|--------|-------|----------|--------|")
|
||||
for name in sorted(stats):
|
||||
s = stats[name]
|
||||
years = ", ".join(f"{y}:{v:.4f}" for y, v in s["byYear"].items()) or "—"
|
||||
|
||||
def _f(v, nd=4):
|
||||
return "—" if v is None else f"{v:.{nd}f}"
|
||||
|
||||
lines.append(f"| {name} | {_f(s['icAll'], 6)} | {_f(s['tAll'])} "
|
||||
f"| {_f(s['positiveRatio'])} | {years} |")
|
||||
```
|
||||
|
||||
(`lines` 为该函数既有的输出行列表变量名,按现场适配。)
|
||||
|
||||
- [ ] **Step 6: 回归+commit**
|
||||
|
||||
Run: `python3 -m pytest tests/factor/ -q`
|
||||
Expected: 全绿
|
||||
|
||||
```bash
|
||||
git add sanguo_factor/monthly_review.py tests/factor/test_monthly_review.py
|
||||
git commit -m "feat(factor): 月度批评报告判定弹药三件——factor_stats 聚合(全期合成t/正IC占比/逐年IC)+md 渲染 [nas]"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 2: detail 端点透传 icStats
|
||||
|
||||
**Files:**
|
||||
- Modify: `sanguo_api/routes_pipeline.py:582`(factor_detail 函数)
|
||||
- Test: `tests/api/test_routes_pipeline_factors.py`(已有文件追加)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: Task 1 报告 JSON 的 `factor_stats` 块;既有 `_monthly_reports() -> list[tuple[host, as_of, doc]]`
|
||||
- Produces: detail 响应 `factor.icStats: {"icAll","tAll","positiveRatio","byYear"} | None`
|
||||
|
||||
- [ ] **Step 1: 写失败测试**
|
||||
|
||||
在 `tests/api/test_routes_pipeline_factors.py` 追加(沿用该文件既有的 client/tmp registry fixture 风格;报告件用 tmp 目录+`SANGUO_FACTOR_MONTHLY_DIR` 环境变量 monkeypatch——若该 env 尚不存在则经 `_monthly_dir` 的真实 env 名适配,见 Step 3 附注):
|
||||
|
||||
```python
|
||||
def test_factor_detail_ic_stats(monkeypatch, tmp_path):
|
||||
"""detail 端点带出 factor_stats(全期弹药),无报告时如实 None."""
|
||||
import json as _json
|
||||
# 注册一个活因子(沿用文件内既有 helper/fixture 注册路径)
|
||||
...
|
||||
rep_dir = tmp_path / "reports" / "factor_monthly"
|
||||
rep_dir.mkdir(parents=True)
|
||||
doc = {"as_of": "2026-09-30",
|
||||
"monthly_points": {"fa_x": [
|
||||
{"month": "2025-09", "ic_mean": 0.04, "t": 2.1, "count": 1},
|
||||
{"month": "2025-10", "ic_mean": 0.06, "t": 2.4, "count": 1},
|
||||
{"month": "2026-09", "ic_mean": 0.02, "t": 1.1, "count": 1}]},
|
||||
"factor_stats": {"fa_x": {"icAll": 0.04, "tAll": 1.7321,
|
||||
"positiveRatio": 1.0,
|
||||
"byYear": {"2025": 0.05, "2026": 0.02}}}}
|
||||
(rep_dir / "nas_2026-09-30.json").write_text(_json.dumps(doc), "utf-8")
|
||||
monkeypatch.setenv("SANGUO_FACTOR_MONTHLY_DIR", str(rep_dir))
|
||||
r = client.get("/pipeline/factors/fa_x/detail")
|
||||
assert r.status_code == 200
|
||||
got = r.json()["factor"]["icStats"]
|
||||
assert got["icAll"] == 0.04 and got["tAll"] == 1.7321
|
||||
assert got["byYear"]["2025"] == 0.05
|
||||
```
|
||||
|
||||
(注册段 `...` 处复用本文件既有测试的注册代码块原样抄——每个测试自含注册是本文件既有惯例。)
|
||||
|
||||
- [ ] **Step 2: 跑测试确认失败**
|
||||
|
||||
Run: `python3 -m pytest tests/api/test_routes_pipeline_factors.py::test_factor_detail_ic_stats -v`
|
||||
Expected: FAIL `KeyError: 'icStats'`(或 assert None 异常)
|
||||
|
||||
- [ ] **Step 3: 实现**
|
||||
|
||||
`factor_detail` 内 `return` 前加:
|
||||
|
||||
```python
|
||||
stats = None
|
||||
reports = _monthly_reports()
|
||||
if reports:
|
||||
stats = (reports[0][2].get("factor_stats") or {}).get(name)
|
||||
```
|
||||
|
||||
返回字面量 `factor` dict 中 `"icRecentT"` 行后加:
|
||||
|
||||
```python
|
||||
"icStats": stats,
|
||||
```
|
||||
|
||||
**附注**:`_monthly_dir()` 若未读 env,则加 `os.environ.get("SANGUO_FACTOR_MONTHLY_DIR", ...)` 前缀(与 `SANGUO_FACTOR_EVAL_DB` 同款覆盖纪律,生产缺省行为不变):
|
||||
|
||||
```python
|
||||
def _monthly_dir() -> str:
|
||||
return os.environ.get("SANGUO_FACTOR_MONTHLY_DIR",
|
||||
os.path.join("reports", "factor_monthly"))
|
||||
```
|
||||
|
||||
- [ ] **Step 4: 跑测试确认通过**
|
||||
|
||||
Run: `python3 -m pytest tests/api/test_routes_pipeline_factors.py -v`
|
||||
Expected: 全 PASS(含既有用例无回归)
|
||||
|
||||
- [ ] **Step 5: Commit**
|
||||
|
||||
```bash
|
||||
git add sanguo_api/routes_pipeline.py tests/api/test_routes_pipeline_factors.py
|
||||
git commit -m "feat(api): 因子详情端点透传 icStats 全期弹药(无报告如实 None)+monthly_dir env 覆盖 [vps]"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 3: expression_match 模块(防换皮算法)
|
||||
|
||||
**Files:**
|
||||
- Create: `sanguo_factor/expression_match.py`
|
||||
- Test: `tests/factor/test_expression_match.py`(新建)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: Python `ast`;表达式方言=factor_guard 白名单(`ts_mean(x, 5)` 函数调用式 + `+ - * /` binop)
|
||||
- Produces:
|
||||
- `similarity(expr_a: str, expr_b: str) -> float | None`(最大公共子树节点数/较小表达式节点数;解析失败 None)
|
||||
- `top_similar(expression: str, candidates: dict[str, str], floor: float = 0.6) -> list[{"name": str, "ratio": float}]`(按 ratio 降序)
|
||||
|
||||
- [ ] **Step 1: 写失败测试**
|
||||
|
||||
```python
|
||||
# tests/factor/test_expression_match.py
|
||||
"""AST 防换皮(spec §4.2 2026-10-10):交换律归一化+最大公共子树占比."""
|
||||
from sanguo_factor.expression_match import similarity, top_similar
|
||||
|
||||
|
||||
def test_identical_and_commutative():
|
||||
assert similarity("close/open", "close/open") == 1.0
|
||||
# 交换律:a+b ≡ b+a(归一化后同构)
|
||||
assert similarity("close + open", "open + close") == 1.0
|
||||
assert similarity("ts_mean(close, 5) * volume",
|
||||
"volume * ts_mean(close, 5)") == 1.0
|
||||
|
||||
|
||||
def test_partial_common_subtree():
|
||||
# 公共子树=ts_mean(close,5)(5节点);分母=较小表达式(7节点)
|
||||
r = similarity("ts_mean(close, 5) / volume",
|
||||
"ts_mean(close, 5) * turnover")
|
||||
assert r is not None and 0.6 < r < 1.0
|
||||
|
||||
|
||||
def test_no_common_and_invalid():
|
||||
assert similarity("close / open", "volume * turnover") == 0.0
|
||||
assert similarity("close +++", "close/open") is None # 解析失败如实 None
|
||||
assert similarity("", "close") is None
|
||||
|
||||
|
||||
def test_top_similar_floor():
|
||||
cands = {"fa_a": "close + open", "fa_b": "open + close",
|
||||
"fa_c": "volume * turnover"}
|
||||
hits = top_similar("close + open", cands, floor=0.6)
|
||||
assert [h["name"] for h in hits] == ["fa_a", "fa_b"]
|
||||
assert all(h["ratio"] >= 0.6 for h in hits)
|
||||
# 空表达式/坏 candidates 不炸
|
||||
assert top_similar("close", {"fa_x": ""}) == []
|
||||
```
|
||||
|
||||
- [ ] **Step 2: 跑测试确认失败**
|
||||
|
||||
Run: `python3 -m pytest tests/factor/test_expression_match.py -v`
|
||||
Expected: FAIL `ModuleNotFoundError: No module named 'sanguo_factor.expression_match'`
|
||||
|
||||
- [ ] **Step 3: 实现模块**
|
||||
|
||||
```python
|
||||
# sanguo_factor/expression_match.py
|
||||
"""表达式结构相似度(AST 防换皮,2026-10-10 spec §4.2).
|
||||
|
||||
QuantaAlpha factor_ast 算法思想(最大公共子树+交换律)的方言内实现:
|
||||
不移植其 qlib 式解析器,直接用 Python ast——与 factor_guard 白名单
|
||||
同方言,注册的新因子表达式必然可解析.判定永远人做:本模块只产提示.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import ast
|
||||
|
||||
_COMMUTATIVE = (ast.Add, ast.Mult)
|
||||
|
||||
|
||||
def _node_size(node: ast.AST) -> int:
|
||||
return 1 + sum(_node_size(c) for c in ast.iter_child_nodes(node))
|
||||
|
||||
|
||||
def _key(node: ast.AST) -> str:
|
||||
"""规范化结构签名:交换律 binop 左右子树排序后拼接."""
|
||||
if isinstance(node, ast.BinOp) and isinstance(node.op, _COMMUTATIVE):
|
||||
lk, rk = _key(node.left), _key(node.right)
|
||||
a, b = sorted((lk, rk))
|
||||
return f"({a}|{b}|{type(node.op).__name__})"
|
||||
if isinstance(node, ast.BinOp):
|
||||
return f"({_key(node.left)}>{_key(node.right)}|{type(node.op).__name__})"
|
||||
if isinstance(node, ast.Call) and isinstance(node.func, ast.Name):
|
||||
args = ",".join(_key(a) for a in node.args)
|
||||
return f"{node.func.id}({args})"
|
||||
if isinstance(node, ast.Name):
|
||||
return f"#{node.id}"
|
||||
if isinstance(node, ast.Constant):
|
||||
return f"#{node.value!r}"
|
||||
return f"?{type(node).__name__}"
|
||||
|
||||
|
||||
def _all_subtree_keys(node: ast.AST) -> dict[str, int]:
|
||||
"""子树签名→节点数(同签名取最大)."""
|
||||
out: dict[str, int] = {}
|
||||
for sub in ast.walk(node):
|
||||
k = _key(sub)
|
||||
out[k] = max(out.get(k, 0), _node_size(sub))
|
||||
return out
|
||||
|
||||
|
||||
def similarity(expr_a: str, expr_b: str) -> float | None:
|
||||
"""最大公共子树节点数 / 较小表达式节点数;任一解析失败=None."""
|
||||
try:
|
||||
ta = ast.parse(expr_a, mode="eval")
|
||||
tb = ast.parse(expr_b, mode="eval")
|
||||
except (SyntaxError, ValueError):
|
||||
return None
|
||||
sa, sb = _node_size(ta), _node_size(tb)
|
||||
if sa == 0 or sb == 0:
|
||||
return None
|
||||
keys_a = _all_subtree_keys(ta)
|
||||
best = 0
|
||||
for sub in ast.walk(tb):
|
||||
k = _key(sub)
|
||||
if k in keys_a:
|
||||
best = max(best, min(keys_a[k], _node_size(sub)))
|
||||
return round(best / min(sa, sb), 4)
|
||||
|
||||
|
||||
def top_similar(expression: str, candidates: dict[str, str],
|
||||
floor: float = 0.6) -> list[dict]:
|
||||
"""对候选池按相似度排序,过滤低于 floor 的(提示非硬拒)."""
|
||||
out: list[dict] = []
|
||||
for name, expr in candidates.items():
|
||||
if not expression or not expr:
|
||||
continue
|
||||
r = similarity(expression, expr)
|
||||
if r is not None and r >= floor:
|
||||
out.append({"name": name, "ratio": r})
|
||||
out.sort(key=lambda h: h["ratio"], reverse=True)
|
||||
return out
|
||||
```
|
||||
|
||||
- [ ] **Step 4: 跑测试确认通过**
|
||||
|
||||
Run: `python3 -m pytest tests/factor/test_expression_match.py -v`
|
||||
Expected: 4 PASS
|
||||
|
||||
- [ ] **Step 5: Commit**
|
||||
|
||||
```bash
|
||||
git add sanguo_factor/expression_match.py tests/factor/test_expression_match.py
|
||||
git commit -m "feat(factor): expression_match 防换皮模块——Python ast 交换律归一化+最大公共子树占比(QuantaAlpha 算法思想方言内实现) [nas]"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 4: decompose 注册链接入 similarity + 端点带出
|
||||
|
||||
**Files:**
|
||||
- Modify: `sanguo_api/routes_pipeline.py:1165` 附近(decompose 注册路径,`vr.upsert_factor(...)` 调用处)+ `factor_detail`/`factors` 端点
|
||||
- Test: `tests/api/test_routes_pipeline_factors.py`(追加)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: Task 3 `top_similar`;registry 条目 `versions[0].params.expression`
|
||||
- Produces: registry 新条目可选键 `similarity: [{"name": str, "ratio": float}]`(仅建条落盘幂等不覆盖——与 description 补丁同纪律);detail 响应 `factor.similarity`
|
||||
|
||||
- [ ] **Step 1: 写失败测试**
|
||||
|
||||
```python
|
||||
def test_decompose_register_similarity_hint(monkeypatch, tmp_path):
|
||||
"""新因子注册时对既有池比对,疑似换皮(≥0.6)存 entry.similarity(提示非硬拒)."""
|
||||
# 复用本文件既有注册 helper:先注册 fa_old(expression="close + open")
|
||||
...
|
||||
# 走 decompose 注册路径注册 fa_new(expression="open + close")
|
||||
# (复用既有 decompose job 测试的注册调用;注册后重读 registry)
|
||||
reg = _load_runtime_registry(tmp_path)
|
||||
entry = reg["factors"]["fa_new"]
|
||||
assert entry.get("similarity") == [{"name": "fa_old", "ratio": 1.0}]
|
||||
|
||||
r = client.get("/pipeline/factors/fa_new/detail")
|
||||
assert r.json()["factor"]["similarity"] == [{"name": "fa_old", "ratio": 1.0}]
|
||||
```
|
||||
|
||||
(两处 `...` 复用本文件既有 decompose/注册测试代码块——file 内已有 decompose job 注册用例,抄其 setup 原样。)
|
||||
|
||||
- [ ] **Step 2: 跑测试确认失败**
|
||||
|
||||
Run: `python3 -m pytest tests/api/test_routes_pipeline_factors.py::test_decompose_register_similarity_hint -v`
|
||||
Expected: FAIL(similarity 键不存在)
|
||||
|
||||
- [ ] **Step 3: 实现(注册点 + 端点)**
|
||||
|
||||
`routes_pipeline.py:1165` 的 `vr.upsert_factor(reg, cand["name"], ...)` 之后、落盘 save 之前加:
|
||||
|
||||
```python
|
||||
from sanguo_factor import expression_match as em
|
||||
_expr = str((cand.get("params") or {}).get("expression") or "")
|
||||
if _expr:
|
||||
_cands = {n: str(((v.get("versions") or [{}])[0]
|
||||
.get("params") or {}).get("expression") or "")
|
||||
for n, v in reg["factors"].items()
|
||||
if n != cand["name"]}
|
||||
_sim = em.top_similar(_expr, _cands)
|
||||
if _sim:
|
||||
reg["factors"][cand["name"]]["similarity"] = _sim
|
||||
```
|
||||
|
||||
(`cand`/`reg` 为该处既有变量名,按现场适配;`upsert_factor` 建新条目后 `reg["factors"][name]` 直接可写。)
|
||||
|
||||
`factor_detail` 返回字面量加:
|
||||
|
||||
```python
|
||||
"similarity": e.get("similarity"),
|
||||
```
|
||||
|
||||
- [ ] **Step 4: 跑测试确认通过+回归**
|
||||
|
||||
Run: `python3 -m pytest tests/api/test_routes_pipeline_factors.py tests/api/ -q`
|
||||
Expected: 全 PASS
|
||||
|
||||
- [ ] **Step 5: Commit**
|
||||
|
||||
```bash
|
||||
git add sanguo_api/routes_pipeline.py tests/api/test_routes_pipeline_factors.py
|
||||
git commit -m "feat(api): decompose 注册时防换皮比对存 similarity 提示(非硬拒,judgement 人做)+detail 带出 [vps]"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 5: 前端展示(弹药卡+换皮徽标)
|
||||
|
||||
**Files:**
|
||||
- Modify: `frontend/src/api/pipeline.ts`(FactorDetail 接口+factor 列表行接口)
|
||||
- Modify: `frontend/src/views/pipeline/FactorDetail.vue`(弹药行+相似提示)
|
||||
- Modify: `frontend/src/views/pipeline/FactorFactory.vue`(列表现有来源列旁加 ⚠ 徽标)
|
||||
- Test: `frontend/src/views/pipeline/FactorDetail.spec.ts`、`FactorFactory.spec.ts`(追加用例)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: detail `factor.icStats`/`factor.similarity`(Task 2/4)
|
||||
- Produces: 展示(无新数据流)
|
||||
|
||||
- [ ] **Step 1: 类型+展示实现**
|
||||
|
||||
`pipeline.ts` FactorDetail 的 factor 接口加:
|
||||
|
||||
```typescript
|
||||
icStats?: { icAll: number | null; tAll: number | null;
|
||||
positiveRatio: number | null; byYear: Record<string, number> } | null
|
||||
similarity?: Array<{ name: string; ratio: number }> | null
|
||||
```
|
||||
|
||||
`FactorDetail.vue` 档案卡 `icRecentT` 行后加(沿用既有 `kv` 行结构):
|
||||
|
||||
```vue
|
||||
<div class="kv"><span class="k">全期弹药</span>
|
||||
<span class="v">
|
||||
<template v-if="detail.factor.icStats && detail.factor.icStats.tAll !== null">
|
||||
全期IC {{ detail.factor.icStats.icAll?.toFixed(4) }} ·
|
||||
合成t {{ detail.factor.icStats.tAll?.toFixed(2) }} ·
|
||||
正IC占比 {{ ((detail.factor.icStats.positiveRatio ?? 0) * 100).toFixed(0) }}% ·
|
||||
逐年 {{ Object.entries(detail.factor.icStats.byYear).map(([y, v]) => `${y}:${v.toFixed(4)}`).join(' ') || '—' }}
|
||||
</template>
|
||||
<template v-else>—(月度点不足或首班未跑)</template>
|
||||
</span></div>
|
||||
<div class="kv" v-if="detail.factor.similarity?.length"><span class="k">疑似换皮</span>
|
||||
<span class="v warn">
|
||||
<span v-for="s in detail.factor.similarity" :key="s.name"
|
||||
class="sim-chip" @click="goFactor(s.name)">
|
||||
{{ s.name }} · {{ (s.ratio * 100).toFixed(0) }}%</span>
|
||||
</span></div>
|
||||
```
|
||||
|
||||
(`goFactor` 复用本组件既有跳转 factor 详情的方法名;`.sim-chip`/`.warn` 样式按本文件既有 style 块补最小两条:cursor:pointer + 底色。)
|
||||
|
||||
`FactorFactory.vue` 表格「来源」列的链接标识后加换皮徽标(行数据经 factors 端点带出 `similarity`——列表端点在 Task 4 一并带出:`factors()` 列表项加 `"similarity": fe.get("similarity")`):
|
||||
|
||||
```vue
|
||||
<el-tooltip v-if="row.similarity?.length"
|
||||
:content="`疑似换皮: ${row.similarity.map(s => s.name).join(', ')}`">
|
||||
<span class="dup-badge" @click="goDetail(row.name)">⚠</span>
|
||||
</el-tooltip>
|
||||
```
|
||||
|
||||
- [ ] **Step 2: 写前端测试(两文件各追加一用例)**
|
||||
|
||||
`FactorDetail.spec.ts`:
|
||||
|
||||
```typescript
|
||||
it('renders ammo stats and similarity chips', () => {
|
||||
const detail = makeDetail()
|
||||
detail.factor.icStats = { icAll: 0.04, tAll: 1.73, positiveRatio: 1.0,
|
||||
byYear: { 2025: 0.05, 2026: 0.02 } }
|
||||
detail.factor.similarity = [{ name: 'fa_old', ratio: 1.0 }]
|
||||
mount(Detail, { props: { name: 'fa_x' } })
|
||||
expect(document.body.textContent).toContain('合成t 1.73')
|
||||
expect(document.body.textContent).toContain('正IC占比 100%')
|
||||
expect(document.body.textContent).toContain('fa_old')
|
||||
})
|
||||
```
|
||||
|
||||
(`makeDetail`/mount 结构抄本文件既有用例;无 icStats 时显示 `—` 的反向用例一并加。)
|
||||
|
||||
`FactorFactory.spec.ts`:
|
||||
|
||||
```typescript
|
||||
it('shows dup badge when similarity present', async () => {
|
||||
const rows = makeRows()
|
||||
rows[0].similarity = [{ name: 'fa_old', ratio: 0.9 }]
|
||||
mount(Factory)
|
||||
await flushPromises()
|
||||
expect(document.querySelector('.dup-badge')).toBeTruthy()
|
||||
})
|
||||
```
|
||||
|
||||
- [ ] **Step 3: 跑测试+类型检查**
|
||||
|
||||
Run: `cd frontend && npx vitest run src/views/pipeline/FactorDetail.spec.ts src/views/pipeline/FactorFactory.spec.ts && npx vue-tsc --noEmit; echo "rc=$?"`
|
||||
Expected: vitest 全 PASS;vue-tsc rc=0(管道退出码显式门控)
|
||||
|
||||
- [ ] **Step 4: Commit**
|
||||
|
||||
```bash
|
||||
git add frontend/src/api/pipeline.ts frontend/src/views/pipeline/FactorDetail.vue frontend/src/views/pipeline/FactorFactory.vue frontend/src/views/pipeline/FactorDetail.spec.ts frontend/src/views/pipeline/FactorFactory.spec.ts
|
||||
git commit -m "feat(frontend): 因子详情全期弹药行+疑似换皮徽标(工厂列表/详情两触点) [vps]"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 6: test 隔离审计 + spec 落账 + 收尾
|
||||
|
||||
**Files:**
|
||||
- Read-only: `sanguo_factor/weight_profiles/quant12_icirfit_v1.json`(fit_window)、`sanguo_factor/composite_weighting.py`、`sanguo_factor/exam_gate.py`(考场窗逻辑)、`sanguo_factor/composite_rolling.py`
|
||||
- Modify: `docs/superpowers/specs/2026-09-22-research-to-trading-pipeline-design.md`(§4.4 test 隔离条款尾部落审计结论)
|
||||
|
||||
**Interfaces:** 无代码——审计+文档。
|
||||
|
||||
- [ ] **Step 1: 审计**
|
||||
|
||||
1. 读 `weight_profiles/quant12_icirfit_v1.json` 的 `fit_window`(ICIR 拟合窗)与 `baseline`;
|
||||
2. 读 `exam_gate.py` 确认 h2 样外考场窗与 fit_window 的关系(考场必须完全后于拟合窗);
|
||||
3. 读 `composite_rolling.py` 确认滚动档案(quant12_roll_20XX.json)的年滚动是否每次只用当年之前数据拟合、之后数据考核;
|
||||
4. 产出三行结论:fit 窗=【】、考场窗=【】、是否重叠=【否/是+风险描述】。
|
||||
|
||||
- [ ] **Step 2: 结论落 spec**
|
||||
|
||||
§4.4 test 隔离条款段尾追加一行:
|
||||
|
||||
```markdown
|
||||
核实结论(2026-10-10 审计):quant12_v2a(quant12_icirfit_v1)fit_window=【实际值】、h2 考场窗=【实际值】——【无重叠,滚动结构天然合规/发现 X,已修或将修】。
|
||||
```
|
||||
|
||||
- [ ] **Step 3: 全量回归**
|
||||
|
||||
Run: `python3 -m pytest tests/factor tests/api -q && cd frontend && npx vitest run src/views/pipeline/ 2>&1 | tail -5`
|
||||
Expected: 全绿
|
||||
|
||||
- [ ] **Step 4: Commit(spec 落账随本 task)**
|
||||
|
||||
```bash
|
||||
git add docs/superpowers/specs/2026-09-22-research-to-trading-pipeline-design.md
|
||||
git commit -m "docs: 流水线 spec 判定弹药/AST防换皮/三路对拍/test隔离四块增补+quant12_v2a 权重窗审计结论 [no-doc]"
|
||||
```
|
||||
|
||||
(spec 增补主体若已在开工前随本地 commit 落盘则本条合并入上一条;`[no-doc]` 用于纯落账——若同 push 含 sanguo_factor 设计变更则改 `[nas]` 并确认 spec 同 push 已更新,档随码走铁律。)
|
||||
|
||||
- [ ] **Step 5: push 纪律(人工闸门)**
|
||||
|
||||
**今晚 21:30 双机日度首班跑完并验证通过后(明早)**,统一 push:
|
||||
|
||||
```bash
|
||||
git log --oneline origin/master..HEAD # 报清单给用户
|
||||
git push origin master # 用户确认后
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Self-Review 结论
|
||||
|
||||
1. **Spec 覆盖**:§4.2 判定弹药三件→Task 1/2/5;§4.2 AST 防换皮→Task 3/4/5;§4.4 test 隔离→Task 6。§4.4 三路对拍(LGBM challenger)**不在本 plan**——班次二单独开计划(spec 已定案)。
|
||||
2. **占位符**:Task 1 Step 5 的 `lines` 变量名、Task 4 的 `cand`/`reg` 变量名、前端 `goFactor`/`makeDetail` 复用点均为「按现场既有名适配」的显式指令(非 TBD),代码块本身完整。
|
||||
3. **类型一致**:`factor_stats` 返回键 `icAll/tAll/positiveRatio/byYear` 在 Task 1(产生)、Task 2(透传)、Task 5(前端 `icStats` 类型)三处一致;`similarity` 的 `{name, ratio}` 在 Task 3/4/5 一致。
|
||||
@@ -0,0 +1,579 @@
|
||||
# LGBM Challenger 三路加权对拍 Implementation Plan(班次二)
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||
|
||||
**Goal:** 同池同窗三路组合信号(等权/ICIR/LightGBM)月度 walk-forward 影子对拍,产出对拍件+特征贡献,回答「ICIR 加权是不是最优」——challenger 永不进生产。
|
||||
|
||||
**Architecture:** 新模块 `sanguo_factor/challenger_lgbm.py`:读月度批已导出的因子值宽表(`monthly_batch --factor-values-out` 产物,每因子一份 datetime×vt_symbol parquet)+ vnpy_db close 构造防缺口 label → 池内三路信号(等权带方向翻转/ICIR 档案权重/LGBM 月度重训)→ 样外月 walk-forward 滚动 → 对拍件 JSON(三路样外 IC/ICIR/分层多空+feature importance+池准入表)。挂 NAS 月度链 stage6(失败不阻链)。
|
||||
|
||||
**Tech Stack:** lightgbm==4.6.0(docker 依赖已在 requirements-docker.txt:84;本机 venv 需装)、pandas/polars 既有栈、pytest。
|
||||
|
||||
## Global Constraints
|
||||
|
||||
- spec=2026-09-22 流水线总设计 §4.4「三路加权对拍(2026-10-10 定案)」+「test 隔离条款」。
|
||||
- **v1 对拍池=quant12_icirfit_v1.sources 的 12 源**(vma_60/alpha16/alpha83/alpha12/alpha2/alpha42/vol_ma5/wvma_20/klow/cord_5/kup/alpha81)——在役晋级池,ICIR 权重档案现成,三路同池才可比;fa_* 新因子扩池挂后续。
|
||||
- **test 隔离铁律(本班直接适用)**:LGBM 训练窗=样外月之前的全部数据,样外月数据绝不进训练;对拍件如实记录切分。
|
||||
- **纪律**:challenger 只落 `reports/factor_monthly/challenger_lgbm/` 子目录(判定层端点只读根层零影响);永不进生产,升级走决议 K 考场+用户口令。
|
||||
- **池贪心准入 v1 形态**:12 源全量进(在役池已过晋级闸),corr 去冗余只**记录**不剔除(`|corr|>0.7` 对标记注在对拍件,剔除动作留人判断——首年观察期)。
|
||||
- label 口径:`close.shift(-2)/close.shift(-1)-1`(T 信号→T+1 收盘可成交→T+2 收盘卖,QuantaAlpha 防缺口同构保守口径);**双侧 CSRankNorm**(特征与 label 截面秩归一 `(rank(pct)-0.5)`,NaN 保持 NaN 不参与当日截面)。
|
||||
- 等权信号=**方向调整后等权**(`direction=='-'` 的源先取负再平均,否则对负向因子不公平);ICIR 同理带 direction;LGBM 不带(模型自学)。
|
||||
- 环境:NAS docker 镜像已含 lightgbm;**本机 venv310 需 `pip install lightgbm==4.6.0`**(dev/test 用);VPS 不跑本链(纯 NAS 影子)→ commit 标签 `[nas]`。
|
||||
- 月度批导出件 12M 窗;walk-forward=前 11M 训练+最后 1M 样外,每月滚动,样外月序列逐月累积。
|
||||
- LGBM 超参(qlib Alpha158 基准起步):`loss=mse, lr=0.1, max_depth=8, num_leaves=210, colsample_bytree=0.8879, subsample=0.8789, lambda_l1=205.6999, lambda_l2=580.9768, min_child_samples=100, feature_fraction_bynode=0.8, seed=42`;early stopping=**无独立 valid 段时用固定 num_boost_round=500**(12M 窗内再切 valid 会挤占训练数据,v1 从简,注记在对拍件)。
|
||||
|
||||
---
|
||||
|
||||
### Task 1: challenger_lgbm 数据层(装载+label+CSRankNorm)
|
||||
|
||||
**Files:**
|
||||
- Create: `sanguo_factor/challenger_lgbm.py`
|
||||
- Test: `tests/factor/test_challenger_lgbm.py`
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `sanguo_factor/weight_profiles/quant12_icirfit_v1.json`(sources 结构 `{name: {direction, fit_icir, weight}}`)
|
||||
- Produces:
|
||||
- `load_pool() -> dict[str, dict]`(12 源档案)
|
||||
- `build_label(close_wide: pd.DataFrame) -> pd.DataFrame`(防缺口 label 宽表)
|
||||
- `cs_rank_norm(df: pd.DataFrame) -> pd.DataFrame`(截面秩归一,NaN 透传)
|
||||
|
||||
- [ ] **Step 1: venv 装 lightgbm**
|
||||
|
||||
```bash
|
||||
venv310/bin/pip install lightgbm==4.6.0
|
||||
venv310/bin/python -c "import lightgbm; print(lightgbm.__version__)"
|
||||
```
|
||||
Expected: `4.6.0`(装不上则报告并停——后续 task 全依赖它)。
|
||||
|
||||
- [ ] **Step 2: 写失败测试**
|
||||
|
||||
```python
|
||||
# tests/factor/test_challenger_lgbm.py
|
||||
"""LGBM challenger 三路对拍(spec §4.4 2026-10-10):数据层."""
|
||||
import json
|
||||
|
||||
import numpy as np
|
||||
import pandas as pd
|
||||
import pytest
|
||||
|
||||
from sanguo_factor.challenger_lgbm import build_label, cs_rank_norm, load_pool
|
||||
|
||||
|
||||
def test_load_pool_twelve_sources():
|
||||
pool = load_pool()
|
||||
assert len(pool) == 12
|
||||
assert pool["vma_60"]["direction"] == "+"
|
||||
assert pool["vol_ma5"]["direction"] == "-"
|
||||
assert 0.0 < pool["alpha16"]["weight"] < 0.2
|
||||
|
||||
|
||||
def test_build_label_gap_proof():
|
||||
idx = pd.date_range("2025-01-01", periods=4, freq="D")
|
||||
close = pd.DataFrame({"a": [10.0, 11.0, 12.0, 13.0], "b": [20.0, 20.0, 21.0, 22.0]}, index=idx)
|
||||
lab = build_label(close)
|
||||
# label[t]=close[t+2]/close[t+1]-1:t0 行=12/11-1
|
||||
assert lab.iloc[0]["a"] == pytest.approx(12.0 / 11.0 - 1)
|
||||
# 末两行无可成交区间=NaN
|
||||
assert lab.iloc[-1].isna().all() and lab.iloc[-2].isna().all()
|
||||
|
||||
|
||||
def test_cs_rank_norm_nan_passthrough_and_centered():
|
||||
df = pd.DataFrame({"a": [1.0, 2.0, 3.0, np.nan],
|
||||
"b": [4.0, 3.0, 2.0, 1.0]})
|
||||
out = cs_rank_norm(df)
|
||||
assert out.loc[2, "a"] == pytest.approx(0.75) # rank pct 1.0 - 0.5
|
||||
assert np.isnan(out.loc[3, "a"]) # NaN 透传不占截面
|
||||
row0 = out.loc[0, ["a", "b"]].tolist()
|
||||
assert row0 == [pytest.approx(-0.25), pytest.approx(0.25)] # 双侧中心化
|
||||
```
|
||||
|
||||
- [ ] **Step 3: 跑测试确认失败**
|
||||
|
||||
Run: `venv310/bin/python -m pytest tests/factor/test_challenger_lgbm.py -v`
|
||||
Expected: FAIL `ModuleNotFoundError`/`ImportError`
|
||||
|
||||
- [ ] **Step 4: 实现数据层**
|
||||
|
||||
```python
|
||||
# sanguo_factor/challenger_lgbm.py
|
||||
"""LGBM challenger 三路加权对拍(2026-10-10 spec §4.4 定案).
|
||||
|
||||
三路=等权(方向调整)/ICIR(quant12_icirfit_v1 档案)/LightGBM(月度重训),
|
||||
同池同窗 walk-forward 影子对拍;challenger 永不进生产,对拍件落
|
||||
reports/factor_monthly/challenger_lgbm/ 子目录(判定层端点只读根层).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
import pandas as pd
|
||||
|
||||
_PROFILE = Path(__file__).parent / "weight_profiles" / "quant12_icirfit_v1.json"
|
||||
|
||||
|
||||
def load_pool() -> dict[str, dict]:
|
||||
"""v1 对拍池=在役 12 源(ICIR 档案现成,三路同池才可比;扩池挂后续)."""
|
||||
with open(_PROFILE, encoding="utf-8") as f:
|
||||
return json.load(f)["sources"]
|
||||
|
||||
|
||||
def build_label(close_wide: pd.DataFrame) -> pd.DataFrame:
|
||||
"""防缺口 label=T+1 收盘→T+2 收盘(信号 T 收盘出,T+1 全天可成交)."""
|
||||
return close_wide.shift(-2) / close_wide.shift(-1) - 1.0
|
||||
|
||||
|
||||
def cs_rank_norm(df: pd.DataFrame) -> pd.DataFrame:
|
||||
"""截面秩归一 (rank_pct-0.5);NaN 透传不占当日截面."""
|
||||
return df.rank(axis=1, pct=True) - 0.5
|
||||
```
|
||||
|
||||
- [ ] **Step 5: 跑测试确认通过并 commit**
|
||||
|
||||
Run: `venv310/bin/python -m pytest tests/factor/test_challenger_lgbm.py -v`
|
||||
Expected: 3 PASS
|
||||
|
||||
```bash
|
||||
git add sanguo_factor/challenger_lgbm.py tests/factor/test_challenger_lgbm.py
|
||||
git commit -m "feat(factor): challenger_lgbm 数据层——在役池档案/防缺口 label/双侧 CSRankNorm [nas]"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 2: 三路信号构造(等权/ICIR/LGBM walk-forward)
|
||||
|
||||
**Files:**
|
||||
- Modify: `sanguo_factor/challenger_lgbm.py`
|
||||
- Test: `tests/factor/test_challenger_lgbm.py`(追加)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: Task 1 三函数;因子值字典 `{name: pd.DataFrame}`(宽表,统一对齐列/索引)
|
||||
- Produces:
|
||||
- `equal_weight_signal(values: dict[str, pd.DataFrame], pool: dict) -> pd.DataFrame`(方向调整等权)
|
||||
- `icir_signal(values: dict, pool: dict) -> pd.DataFrame`(档案权重加权)
|
||||
- `lgbm_walk_forward(values: dict, label: pd.DataFrame, last_month: str, params: dict | None = None) -> tuple[pd.DataFrame, dict]`(返回样外月预测+feature importance;训练窗=样外月之前全部)
|
||||
|
||||
- [ ] **Step 1: 写失败测试**
|
||||
|
||||
```python
|
||||
def _toy_values(idx, cols):
|
||||
rng = np.random.default_rng(7)
|
||||
return {n: pd.DataFrame(rng.normal(size=(len(idx), len(cols))), index=idx, columns=cols)
|
||||
for n in ("f_pos", "f_neg")}
|
||||
|
||||
|
||||
def test_equal_weight_direction_adjusted():
|
||||
idx = pd.date_range("2025-06-01", periods=3, freq="D")
|
||||
cols = ["a", "b"]
|
||||
values = {"f_pos": pd.DataFrame(1.0, index=idx, columns=cols),
|
||||
"f_neg": pd.DataFrame(1.0, index=idx, columns=cols)}
|
||||
pool = {"f_pos": {"direction": "+"}, "f_neg": {"direction": "-"}}
|
||||
sig = equal_weight_signal(values, pool)
|
||||
# +1 与 -1 等权平均=0
|
||||
assert (sig == 0.0).all().all()
|
||||
|
||||
|
||||
def test_icir_signal_uses_archive_weights():
|
||||
idx = pd.date_range("2025-06-01", periods=2, freq="D")
|
||||
cols = ["a"]
|
||||
values = {"f_pos": pd.DataFrame(2.0, index=idx, columns=cols),
|
||||
"f_neg": pd.DataFrame(2.0, index=idx, columns=cols)}
|
||||
pool = {"f_pos": {"direction": "+", "weight": 0.75},
|
||||
"f_neg": {"direction": "-", "weight": 0.25}}
|
||||
sig = icir_signal(values, pool)
|
||||
assert (sig == pytest.approx(2.0 * (0.75 - 0.25))).all().all()
|
||||
|
||||
|
||||
def test_lgbm_walk_forward_holdout_isolated():
|
||||
"""样外月绝不进训练(test 隔离铁律):给训练月与样外月截然不同的
|
||||
因子-收益关系,样外预测应反映训练期学到的关系而非记忆样外."""
|
||||
import pandas as pd
|
||||
from sanguo_factor.challenger_lgbm import build_label, lgbm_walk_forward
|
||||
idx = pd.date_range("2025-01-01", "2025-03-31", freq="B")
|
||||
cols = [f"s{i}" for i in range(5)]
|
||||
rng = np.random.default_rng(42)
|
||||
values = {"s0": pd.DataFrame(rng.normal(size=(len(idx), 5)), index=idx, columns=cols)}
|
||||
label = build_label(pd.DataFrame(100 + rng.normal(scale=0.5, size=(len(idx), 5)),
|
||||
index=idx, columns=cols))
|
||||
oos = idx[idx >= "2025-03-01"]
|
||||
pred, imp = lgbm_walk_forward(values, label, last_month="2025-03")
|
||||
assert list(pred.index) == [d for d in oos if d in label.index and label.loc[d].notna().any()]
|
||||
assert set(imp.keys()) == {"s0"} and imp["s0"] > 0
|
||||
```
|
||||
|
||||
- [ ] **Step 2: 跑测试确认失败** — `ImportError`(三函数未定义)
|
||||
|
||||
- [ ] **Step 3: 实现**
|
||||
|
||||
```python
|
||||
LGBM_PARAMS = { # qlib Alpha158 基准超参起步(spec §4.4;无独立 valid 段,v1 固定轮数)
|
||||
"objective": "mse", "learning_rate": 0.1, "max_depth": 8, "num_leaves": 210,
|
||||
"colsample_bytree": 0.8879, "subsample": 0.8789,
|
||||
"lambda_l1": 205.6999, "lambda_l2": 580.9768,
|
||||
"min_child_samples": 100, "feature_fraction_bynode": 0.8,
|
||||
"seed": 42, "num_threads": 4, "verbose": -1,
|
||||
}
|
||||
NUM_BOOST_ROUND = 500
|
||||
|
||||
|
||||
def _aligned(values: dict[str, pd.DataFrame]) -> pd.DataFrame:
|
||||
"""多因子宽表纵向拼接成特征长表(index 对齐,缺失源列 NaN)."""
|
||||
return pd.concat({n: cs_rank_norm(v) for n, v in values.items()}, axis=1)
|
||||
|
||||
|
||||
def _dir_sign(pool: dict, name: str) -> float:
|
||||
return -1.0 if pool.get(name, {}).get("direction") == "-" else 1.0
|
||||
|
||||
|
||||
def equal_weight_signal(values: dict[str, pd.DataFrame], pool: dict) -> pd.DataFrame:
|
||||
stack = pd.concat([_dir_sign(pool, n) * cs_rank_norm(v) for n, v in values.items()])
|
||||
return stack.groupby(level=0).mean()
|
||||
|
||||
|
||||
def icir_signal(values: dict[str, pd.DataFrame], pool: dict) -> pd.DataFrame:
|
||||
total = sum(pool[n].get("weight", 0.0) for n in values)
|
||||
if total <= 0:
|
||||
raise ValueError("ICIR 权重和为零,档案异常")
|
||||
out = None
|
||||
for n, v in values.items():
|
||||
w = pool[n].get("weight", 0.0) * _dir_sign(pool, n)
|
||||
part = w * cs_rank_norm(v)
|
||||
out = part if out is None else out.add(part, fill_value=0.0)
|
||||
return out / total
|
||||
|
||||
|
||||
def lgbm_walk_forward(values: dict[str, pd.DataFrame], label: pd.DataFrame,
|
||||
last_month: str, params: dict | None = None) -> tuple[pd.DataFrame, dict]:
|
||||
"""训练窗=last_month 之前全部;样外=last_month 当月(test 隔离铁律).
|
||||
|
||||
返回 (样外日×股票预测宽表, {feature: gain}).特征=CSRankNorm 后各源,
|
||||
label=CSRankNorm 后防缺口收益;日频截面样本(日期,股票)平铺训练.
|
||||
"""
|
||||
import lightgbm as lgb
|
||||
|
||||
feat = _aligned(values)
|
||||
lab = cs_rank_norm(label)
|
||||
common = feat.index.intersection(lab.index)
|
||||
feat, lab = feat.loc[common], lab.loc[common]
|
||||
|
||||
oos_mask = feat.index.strftime("%Y-%m") == last_month
|
||||
train_mask = ~oos_mask
|
||||
X_tr = feat[train_mask].stack(future_stack=True).reset_index()
|
||||
X_tr.columns = ["datetime", "vt_symbol", *feat.columns.levels[0]]
|
||||
y_df = lab.stack(future_stack=True).rename("y").reset_index()
|
||||
tr = X_tr.merge(y_df, on=["datetime", "vt_symbol"]).dropna()
|
||||
X = tr[list(feat.columns.levels[0])]
|
||||
model = lgb.train(params or LGBM_PARAMS, lgb.Dataset(X, label=tr["y"]),
|
||||
num_boost_round=NUM_BOOST_ROUND)
|
||||
imp = dict(zip(X.columns, model.feature_importance("gain").tolist()))
|
||||
|
||||
X_oos = feat[oos_mask].stack(future_stack=True).reset_index()
|
||||
X_oos.columns = X_tr.columns
|
||||
X_oos = X_oos.dropna(subset=list(feat.columns.levels[0]))
|
||||
if X_oos.empty:
|
||||
return pd.DataFrame(), imp
|
||||
preds = model.predict(X_oos[list(feat.columns.levels[0])])
|
||||
out = X_oos[["datetime", "vt_symbol"]].assign(p=preds)
|
||||
return out.pivot(index="datetime", columns="vt_symbol", values="p"), imp
|
||||
```
|
||||
|
||||
(`stack(future_stack=True)` 为 pandas≥2.1 语义——本仓 pandas 版本若 <2.1 改 `stack(dropna=True)`,跑 Step 4 时确认。)
|
||||
|
||||
- [ ] **Step 4: 跑测试确认通过并 commit**
|
||||
|
||||
Run: `venv310/bin/python -m pytest tests/factor/test_challenger_lgbm.py -v`
|
||||
Expected: 6 PASS
|
||||
|
||||
```bash
|
||||
git add sanguo_factor/challenger_lgbm.py tests/factor/test_challenger_lgbm.py
|
||||
git commit -m "feat(factor): challenger 三路信号——方向调整等权/ICIR 档案加权/LGBM walk-forward(样外隔离铁律) [nas]"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 3: 对拍评估与产物件
|
||||
|
||||
**Files:**
|
||||
- Modify: `sanguo_factor/challenger_lgbm.py`
|
||||
- Test: `tests/factor/test_challenger_lgbm.py`(追加)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: Task 2 三信号+`build_label`
|
||||
- Produces:
|
||||
- `score_signal(signal: pd.DataFrame, label: pd.DataFrame) -> dict`(`{"ic_mean","icir","q5q1","days"}`——日 IC 均值/ICIR/分层多空累计/有效天数)
|
||||
- `run_challenge(values_dir: str, vnpy_db: str, as_of: str, out_dir: str) -> str`(主入口:读月度批导出宽表+拉 close→三路→对拍件 JSON 落 `out_dir/challenger_lgbm/{host}_{as_of}.json`,返回件路径)
|
||||
|
||||
- [ ] **Step 1: 写失败测试**
|
||||
|
||||
```python
|
||||
def test_score_signal_perfect_and_flat():
|
||||
idx = pd.date_range("2025-03-03", periods=10, freq="B")
|
||||
cols = ["a", "b", "c"]
|
||||
sig = pd.DataFrame(np.linspace(-1, 1, 30).reshape(10, 3), index=idx, columns=cols)
|
||||
lab = sig * 1.0 # 完美信号
|
||||
s = score_signal(sig, lab)
|
||||
assert s["ic_mean"] == pytest.approx(1.0, abs=1e-6)
|
||||
assert s["q5q1"] > 0
|
||||
flat = pd.DataFrame(0.0, index=idx, columns=cols)
|
||||
s0 = score_signal(flat, lab)
|
||||
assert s0["days"] == 10
|
||||
|
||||
|
||||
def test_run_challenge_end_to_end(tmp_path, monkeypatch):
|
||||
"""端到端:合成 3 因子×40 日数据落宽表 parquet+合成 close db 太重——
|
||||
本用例走 values_dir 真文件+vnpy_db 用 monkeypatch 替换拉取函数."""
|
||||
from sanguo_factor import challenger_lgbm as cl
|
||||
idx = pd.date_range("2025-01-01", "2025-03-31", freq="B")
|
||||
cols = ["a", "b"]
|
||||
rng = np.random.default_rng(3)
|
||||
values = {n: pd.DataFrame(rng.normal(size=(len(idx), 2)), index=idx, columns=cols)
|
||||
for n in ("s0", "s1")}
|
||||
vdir = tmp_path / "factor_values"
|
||||
vdir.mkdir()
|
||||
for n, v in values.items():
|
||||
v.to_parquet(vdir / f"{n}.parquet")
|
||||
close = pd.DataFrame(100 + np.cumsum(rng.normal(scale=0.4, size=(len(idx), 2)), axis=0),
|
||||
index=idx, columns=cols)
|
||||
monkeypatch.setattr(cl, "_load_close_wide", lambda db, cols_ref: close)
|
||||
out = cl.run_challenge(str(vdir), "fake.db", "2025-03-31", str(tmp_path))
|
||||
doc = json.loads(Path(out).read_text())
|
||||
assert doc["as_of"] == "2025-03-31"
|
||||
assert set(doc["signals"]) == {"equal", "icir", "lgbm"}
|
||||
for k, v in doc["signals"].items():
|
||||
assert {"ic_mean", "icir", "q5q1", "days"} <= set(v)
|
||||
assert doc["pool"]["names"] == ["s0", "s1"]
|
||||
assert len(doc["feature_importance"]) == 2
|
||||
assert doc["split"]["oos_month"] == "2025-03"
|
||||
```
|
||||
|
||||
- [ ] **Step 2: 跑测试确认失败**
|
||||
|
||||
- [ ] **Step 3: 实现**
|
||||
|
||||
```python
|
||||
def _load_close_wide(vnpy_db: str, columns_ref: pd.Index) -> pd.DataFrame:
|
||||
"""按因子宽表列(股票)拉收盘价.生产实现走 load_universe_bars 轻量列."""
|
||||
from sanguo_data.config import load_config, find_config_path
|
||||
from .batch_eval import load_universe_bars
|
||||
cfg = load_config(find_config_path())
|
||||
bars = load_universe_bars(vnpy_db or cfg.data_paths["vnpy_db"],
|
||||
symbols=list(columns_ref), limit=None)
|
||||
_ = bars.select(["datetime", "vt_symbol", "close"]).to_pandas()
|
||||
wide = _.pivot(index="datetime", columns="vt_symbol", values="close").sort_index()
|
||||
wide.index = pd.to_datetime(wide.index)
|
||||
return wide
|
||||
|
||||
|
||||
def score_signal(signal: pd.DataFrame, label: pd.DataFrame) -> dict:
|
||||
"""样外评分:日 IC 均值/ICIR/五分位多空累计/有效天数."""
|
||||
common = signal.index.intersection(label.index)
|
||||
ics = []
|
||||
ls_rets = []
|
||||
for d in common:
|
||||
s, y = signal.loc[d], label.loc[d]
|
||||
pair = pd.concat([s, y], axis=1, keys=["s", "y"]).dropna()
|
||||
if len(pair) < 5:
|
||||
continue
|
||||
ics.append(pair["s"].corr(pair["y"], method="spearman"))
|
||||
q = pair["s"].quantile([0.2, 0.8])
|
||||
lo, hi = pair[pair["s"] <= q[0.2]]["y"].mean(), pair[pair["s"] >= q[0.8]]["y"].mean()
|
||||
ls_rets.append((hi - lo) if (lo is not None and hi is not None) else 0.0)
|
||||
if not ics:
|
||||
return {"ic_mean": None, "icir": None, "q5q1": None, "days": 0}
|
||||
ser = pd.Series(ics).dropna()
|
||||
icir = (ser.mean() / ser.std()) if len(ser) > 1 and ser.std() > 0 else None
|
||||
return {"ic_mean": round(float(ser.mean()), 6),
|
||||
"icir": round(float(icir), 6) if icir is not None else None,
|
||||
"q5q1": round(float(sum(ls_rets)), 6), "days": len(ser)}
|
||||
|
||||
|
||||
def run_challenge(values_dir: str, vnpy_db: str, as_of: str, out_dir: str,
|
||||
host: str = "nas") -> str:
|
||||
"""月度链 stage6 入口:读导出宽表→三路→对拍件(append-only 子目录)."""
|
||||
pool = load_pool()
|
||||
names = sorted(pool)
|
||||
values = {n: pd.read_parquet(os.path.join(values_dir, f"{n}.parquet"))
|
||||
for n in names if os.path.exists(os.path.join(values_dir, f"{n}.parquet"))}
|
||||
if not values:
|
||||
raise FileNotFoundError(f"因子宽表目录无池内因子: {values_dir}")
|
||||
cols_ref = next(iter(values.values())).columns
|
||||
close = _load_close_wide(vnpy_db, cols_ref)
|
||||
label = build_label(close)
|
||||
last_month = as_of[:7]
|
||||
|
||||
signals = {"equal": equal_weight_signal(values, pool),
|
||||
"icir": icir_signal(values, pool)}
|
||||
pred, imp = lgbm_walk_forward(values, label, last_month)
|
||||
if not pred.empty:
|
||||
signals["lgbm"] = pred
|
||||
|
||||
oos_label = label[label.index.strftime("%Y-%m") == last_month]
|
||||
scored = {k: score_signal(v, oos_label) for k, v in signals.items()}
|
||||
|
||||
# 池冗余注记:|corr|>0.7 只记录不剔除(首年观察期,剔除留人)
|
||||
flat = pd.concat({n: cs_rank_norm(v) for n, v in values.items()})
|
||||
daily_corr = flat.groupby(level=0).apply(
|
||||
lambda g: g.T.corr(method="spearman") if g.shape[0] > 1 else None)
|
||||
flagged = []
|
||||
means = daily_corr.groupby(level=1).mean() if daily_corr is not None else None
|
||||
if means is not None:
|
||||
for n1 in means.index:
|
||||
for n2 in means.columns:
|
||||
if n1 < n2 and abs(means.loc[n1, n2]) > 0.7:
|
||||
flagged.append({"a": n1, "b": n2,
|
||||
"corr": round(float(means.loc[n1, n2]), 4)})
|
||||
|
||||
import platform
|
||||
doc = {"as_of": as_of, "generated_at": pd.Timestamp.now().isoformat(),
|
||||
"host": host, "pool": {"names": sorted(values), "size": len(values)},
|
||||
"split": {"oos_month": last_month,
|
||||
"note": "训练=样外月前全部(test 隔离铁律);无独立 valid,固定轮数"},
|
||||
"signals": scored, "feature_importance": imp,
|
||||
"redundancy_flagged": flagged,
|
||||
"lgbm_params": {**LGBM_PARAMS, "num_boost_round": NUM_BOOST_ROUND}}
|
||||
sub = os.path.join(out_dir, "challenger_lgbm")
|
||||
os.makedirs(sub, exist_ok=True)
|
||||
path = os.path.join(sub, f"{host}_{as_of}.json")
|
||||
with open(path, "w", encoding="utf-8") as f:
|
||||
json.dump(doc, f, ensure_ascii=False, indent=2)
|
||||
return path
|
||||
```
|
||||
|
||||
- [ ] **Step 4: 跑测试确认通过并 commit**
|
||||
|
||||
Run: `venv310/bin/python -m pytest tests/factor/test_challenger_lgbm.py -v`
|
||||
Expected: 8 PASS
|
||||
|
||||
```bash
|
||||
git add sanguo_factor/challenger_lgbm.py tests/factor/test_challenger_lgbm.py
|
||||
git commit -m "feat(factor): challenger 对拍评估——三路样外 IC/ICIR/分层多空+特征贡献+池冗余注记件 [nas]"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 4: CLI 入口+月度链 stage6 接线
|
||||
|
||||
**Files:**
|
||||
- Modify: `sanguo_factor/challenger_lgbm.py`(main)
|
||||
- Modify: `scripts/nas_sync/run_factor_monthly_standalone.sh`(stage5 replay 后加 stage6)
|
||||
- Test: `tests/factor/test_challenger_lgbm.py`(CLI 冒烟一用例)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: Task 3 `run_challenge`;月度批 `--factor-values-out` 导出目录(NAS wrapper 需确认 monthly_batch CLI 透传参数名——读 monthly_batch main 后按实际名接)
|
||||
- Produces: `python -m sanguo_factor.challenger_lgbm --values-dir ... --as-of ... --out-dir ... [--vnpy-db ...] [--host nas]`;stage6 产物 `reports/factor_monthly/challenger_lgbm/nas_{as_of}.json`
|
||||
|
||||
- [ ] **Step 1: 写 CLI 冒烟失败测试**
|
||||
|
||||
```python
|
||||
def test_cli_smoke(tmp_path, monkeypatch, capsys):
|
||||
from sanguo_factor import challenger_lgbm as cl
|
||||
idx = pd.date_range("2025-01-01", "2025-03-31", freq="B")
|
||||
close = pd.DataFrame(100 + np.cumsum(np.random.default_rng(1).normal(0.4, size=(len(idx), 2)), axis=0),
|
||||
index=idx, columns=["a", "b"])
|
||||
monkeypatch.setattr(cl, "_load_close_wide", lambda db, c: close)
|
||||
vdir = tmp_path / "vals"; vdir.mkdir()
|
||||
for n in ("s0", "s1"):
|
||||
pd.DataFrame(np.random.default_rng(2).normal(size=(len(idx), 2)),
|
||||
index=idx, columns=["a", "b"]).to_parquet(vdir / f"{n}.parquet")
|
||||
rc = cl.main(["--values-dir", str(vdir), "--as-of", "2025-03-31",
|
||||
"--out-dir", str(tmp_path), "--vnpy-db", "fake.db"])
|
||||
assert rc == 0
|
||||
assert (tmp_path / "challenger_lgbm" / "nas_2025-03-31.json").exists()
|
||||
```
|
||||
|
||||
- [ ] **Step 2: 确认失败 → 实现 main + 接线**
|
||||
|
||||
main(模块尾):
|
||||
|
||||
```python
|
||||
def main(argv=None):
|
||||
import argparse
|
||||
ap = argparse.ArgumentParser(description="LGBM challenger 三路对拍(月度链 stage6)")
|
||||
ap.add_argument("--values-dir", required=True,
|
||||
help="月度批 --factor-values-out 导出目录(每因子一 parquet)")
|
||||
ap.add_argument("--as-of", required=True)
|
||||
ap.add_argument("--out-dir", required=True, help="报告根(factor_monthly)")
|
||||
ap.add_argument("--vnpy-db", default=None)
|
||||
ap.add_argument("--host", default="nas")
|
||||
a = ap.parse_args(argv)
|
||||
try:
|
||||
path = run_challenge(a.values_dir, a.vnpy_db, a.as_of, a.out_dir, host=a.host)
|
||||
except FileNotFoundError as e:
|
||||
print(f"[challenger] skip: {e}")
|
||||
return 0 # 无宽表=非错误(月度批未开导出),不阻链
|
||||
print(f"[challenger] 对拍件: {path}")
|
||||
return 0
|
||||
```
|
||||
|
||||
**wrapper 接线**(`run_factor_monthly_standalone.sh` stage5 段后,最终 rc 聚合前):
|
||||
|
||||
```bash
|
||||
# stage6 challenger 三路对拍(影子实验,失败不阻月度链;产物落子目录
|
||||
# 判定层端点只读根层零影响)
|
||||
"$DOCKER" run --rm --name sanguo-factor-chal --user "${UID_ADMIN}:${GID_ADMIN}" \
|
||||
--group-add "${GID_ADMINS}" --no-healthcheck --entrypoint python \
|
||||
-e HOME=/tmp -e MPLCONFIGDIR=/tmp/mpl \
|
||||
-v /volume1/stock:/volume1/stock \
|
||||
-v "$APP":/app:ro \
|
||||
-w /app \
|
||||
sanguo_vnpy_v2:lock-aligned \
|
||||
-m sanguo_factor.challenger_lgbm \
|
||||
--values-dir "$BASE/reports/factor_values" \
|
||||
--as-of "$AS_OF" --out-dir "$OUT_DIR" --host nas
|
||||
rc_chal=$?
|
||||
echo "=== $(date '+%F %T') stage6 challenger exit=$rc_chal (影子,不进 final) ==="
|
||||
```
|
||||
|
||||
(`$BASE/reports/factor_values` 为月度批 `--factor-values-out` 实际目录——**执行时先读 `monthly_batch.py` CLI 与 wrapper 现状确认导出是否已开+目录名**;若月度批尚未传导出参数,本 task 同时给 wrapper 的 stage1 命令加透传并验证 monthly_batch CLI 支持该 flag,不支持则本 task 给 monthly_batch 补 `--factor-values-out` CLI 参数透传到 run_batch_eval——已存在函数参数,只补 argparse。)
|
||||
|
||||
- [ ] **Step 3: 跑测试+本地冒烟并 commit**
|
||||
|
||||
Run: `venv310/bin/python -m pytest tests/factor/test_challenger_lgbm.py -v && bash -n scripts/nas_sync/run_factor_monthly_standalone.sh`
|
||||
Expected: 9 PASS + bash -n 语法过
|
||||
|
||||
```bash
|
||||
git add sanguo_factor/challenger_lgbm.py scripts/nas_sync/run_factor_monthly_standalone.sh tests/factor/test_challenger_lgbm.py
|
||||
git commit -m "feat(factor): challenger CLI+月度链 stage6 接线(影子失败不阻链;月度批因子宽表导出透传) [nas]"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 5: spec 落账+全量回归+真数据首跑预约
|
||||
|
||||
**Files:**
|
||||
- Modify: `docs/superpowers/specs/2026-09-22-research-to-trading-pipeline-design.md`(§4.4 三路对拍块补「班次二落地」一行)
|
||||
- Test: 全量回归
|
||||
|
||||
- [ ] **Step 1: spec 落账**
|
||||
|
||||
§4.4 三路对拍块尾追加:
|
||||
|
||||
```markdown
|
||||
- **班次二落地(2026-10-10)**:`sanguo_factor/challenger_lgbm.py`——v1 池=quant12 在役 12 源(扩池挂后续);三路=方向调整等权/ICIR 档案加权/LGBM(qlib 基准超参,固定 500 轮——无独立 valid 段不早停,注记在对拍件);walk-forward=前 11M 训练+末 1M 样外(test 隔离铁律:样外月绝不进训练);对拍件=`challenger_lgbm/{host}_{as_of}.json`(三路样外 IC/ICIR/Q5-Q1 多空+feature importance+池冗余注记——`|corr|>0.7` 只记录不剔除,首年观察期剔除留人);挂月度链 stage6(失败不阻链);判读需累积≥2-3 个样外月。
|
||||
```
|
||||
|
||||
- [ ] **Step 2: 全量回归**
|
||||
|
||||
Run: `venv310/bin/python -m pytest tests/factor tests/api -q && cd frontend && npx vitest run src/views/pipeline/ 2>&1 | tail -3`
|
||||
Expected: 全绿(challenger 新 9 用例含其中)
|
||||
|
||||
- [ ] **Step 3: commit**
|
||||
|
||||
```bash
|
||||
git add docs/superpowers/specs/2026-09-22-research-to-trading-pipeline-design.md
|
||||
git commit -m "docs: 三路对拍班次二落地落账 spec §4.4 [no-doc]"
|
||||
```
|
||||
|
||||
(若 Task 4 同 push 含 sanguo_factor 设计变更,spec 同 push 已更新则本条合并标签纪律照旧。)
|
||||
|
||||
- [ ] **Step 4: 真数据首跑预约(不跑)**
|
||||
|
||||
真数据首跑=**11-01 月度链首班自动带 stage6**(NAS wrapper 生效位需 `scp -O scripts/nas_sync/run_factor_monthly_standalone.sh sanguo-nas:/volume1/stock/sanguo_vnpy_v2/run_factor_monthly.sh` 同步——**push 过 nas-verify 后执行**,与既往生效位纪律一致)。本班不本地跑真数据(vnpy_db 全量在 NAS)。
|
||||
|
||||
---
|
||||
|
||||
## Self-Review 结论
|
||||
|
||||
1. **Spec 覆盖**:§4.4 三路(等权 Task 2/ICIR Task 2/LGBM Task 2)、月度重训(walk-forward Task 2)、CSRankNorm 双侧(Task 1)、防缺口 label(Task 1)、池贪心准入 v1=全量+冗余注记(Task 3,首年观察期口径已在 Global Constraints 声明)、feature importance(Task 2/3)、challenger 子目录隔离(Task 3)、影子纪律(Task 4 stage6 不阻链)——全覆盖。**spec 的「贪心准入」v1 简化为记录不剔除**(在役池已过晋级闸,重复剔除是减法实验,留观察期后决定)——与 spec「起步=晋级因子全量不设限」一致。
|
||||
2. **占位符**:无 TBD;Task 4 的「执行时确认 factor-values-out 目录名」是显式适配指令(月度批 CLI 需现场核对)。
|
||||
3. **类型一致**:`cs_rank_norm/build_label/load_pool`(Task 1)→三信号(Task 2)→`score_signal/run_challenge`(Task 3)→`main`(Task 4)签名逐级消费一致;对拍件 JSON 键 `signals/pool/split/feature_importance/redundancy_flagged/lgbm_params` 在 Task 3 产生、测试断言同键。
|
||||
@@ -75,6 +75,17 @@ F 环是「管道」与「飞轮」的分界:归因发现→可能=衰减触
|
||||
- **出生 description 链路(2026-10-09 补丁,用户反馈「光看名字看不明白」)**:分解器输出 schema 四字段(name/expression/**description**/justification)——description=一句话说因子本身是什么(LLM 出生时生成,≤120 字),justification=如何承载假设(既有语义不变);注册时落 registry **条目级**(与 hypothesis/origin 同级,仅建条落盘幂等不覆盖,换描述走升版本);`/pipeline/factors` 与 detail 端点带出;工厂看板因子名下方副行展示+详情页出生档案行。存量因子无 description 显示空,不做迁移(新出生自动有)。
|
||||
- **详情页送墓园触点(2026-10-09 补丁)**:因子详情页危险区加「送墓园」——状态机本就允许任何活着状态→graveyard(incubating/assessable/promoted/decaying/retired),verdict 端点+判死回写卡片全现成,本补丁只补 UI 入口(此前唯一入口=晋级评审页,只收 assessable 且 t≥2,incubating 因子无处可毙);**note 死因必填**(墓园档案价值所在,弹窗输入);promoted/decaying 态=重确认档(生产在用、合成层钉着版本,文案引导「建议先退役」),incubating/assessable=普通确认档;墓园因子不出工厂板(卡片因子清单 registry 反查自动消失)。**不做物理删除**(Linus 三问不通过:名字释放可绕[决议 D 换假设开新因子]、墓园除噪场景为空、三处置动作概念负担大于收益——墓园=有档案的死亡防重复造轮,足够)。
|
||||
- **因子墓园可见性(2026-10-09 深夜补丁,用户验收「送进墓园去哪看」暴露)**:判死因子的回写链只覆盖 `hyp-` 新卡(H- 老编号静默跳过),且 registry 的 graveyard 档案(name+cause)原本无任何 UI 视图——判死即「不可见」。补:`GET /pipeline/factors/graveyard`(registry graveyard 条目带 cause/hypothesis/origin)+假设池页墓园 tab 加「☠ 判死因子」区(名称划线+死因+无卡兜底文案)——**因子墓园唯一视图**(死假设卡列表仍在同 tab 上方,卡片/因子两个维度并列);夜班批自动判死同走此视图。
|
||||
- **判定弹药三件(2026-10-10 补丁,QuantaAlpha 论文口径借鉴;业务问题=月度批评判定人信息缺口——只有近窗 t,分不清「废了」vs「挨风格打」vs「靠几天撑均值」)**:
|
||||
- **全期 t=出生以来累积均值**:factors 端点/详情页的全史 t 从如实 null 升级为 monthly_points 序列累积均值——数据已有纯聚合,不建出厂全窗批(12M 滚动已覆盖近态,全窗批成本高收益低挂可选 backlog);
|
||||
- **t 统计量+正 IC 天数占比**进月度批评报告与因子详情(t=mean/(std/√n)、正 IC 占比=正点数/总点数;QuantaAlpha 论文 A.5 同款报告纪律,其开源代码亦无——抄论文不抄代码);
|
||||
- **逐年 IC**:报告/工厂看板按年聚合(月度点 groupby year;日度化后点更密判读力自然上来);
|
||||
- **小样本诚实注记**:月度点仅 12 个时 t 判读弱(95% 置信需 t>2.2,小样本难达)——本弹药与日度化配套,点数上量后判读力自然上来;判定动作不因此自动化(宪法:阈值叫人、处置人选)。
|
||||
- **AST 防换皮(2026-10-10 补丁,QuantaAlpha `factors/coder/factor_ast.py` 算法思想方言内重实现;业务问题=换皮因子虚增池数/稀释合成层权重/重复评估)**:
|
||||
- 算法思想方言内重实现(120 行,Python ast,零新依赖):表达式→语法树→最大公共子树匹配,含交换律识别(`a+b`≡`b+a`)+UnaryOp/Call keywords 显式编码;**软提示占比口径(公共子树节点数/较小表达式节点数)与 factor_guard 硬门绝对节点数口径(`_DUP_SUBTREE_MIN=8`)相互独立,不抽公共 canonical**;
|
||||
- 接入点=版本注册表:**注册新因子时**(仅建条落盘,幂等重注册不重写)对既有因子池做结构比对,高相似=**提示非硬拒**(「疑似换皮:与 fXX 最大公共子树占比 X%」,registry 每条存前 5,人决定收不收)——judgement 永远人做,硬拒违反人工卡点宪法;
|
||||
- 方言适配:直接用 Python ast(与 factor_guard 白名单同方言,注册表达式必然可解析),不移植其 qlib 式解析器;
|
||||
- 工厂看板逐年展示落点=列表 `icFullT`(全期合成 t,月报 factor_stats.tAll)+详情 `icStats.byYear`(逐年)——2026-10-10 已兑现;
|
||||
- 不抄 RD-Agent 的 prompt 软门(「请勿生成重复因子」——LLM 自觉防不住换皮,业界反例)。
|
||||
|
||||
### §4.3 晋级闸门
|
||||
|
||||
@@ -83,6 +94,15 @@ F 环是「管道」与「飞轮」的分界:归因发现→可能=衰减触
|
||||
### §4.4 合成层
|
||||
|
||||
- **等权=每个新合成层的基线**(composite_fund8 现状);**权重进化走「考场晋级+版本化」**(决议 K,09-24 辩论裁决):ICIR 拟合形态=从等权基线在 h2 样外考场打赢在位者后以新版本上位(先例 quant12 v1 等权→v2a ICIR,09-11 收官),护栏=非负+单源≤2×等权+年频拟合+0.7 族压缩前置(均已测试锁死),档案铁律=新档案新名字+reason 必填+生产冠军定义永不改。P3 归因贡献类权重(第三形态)攒数据后再议。
|
||||
- **三路加权对拍(2026-10-10 定案;业务问题=「ICIR 加权是不是最优」无对照永拍脑袋——组合因子的核心要素是加权方式,用实验攒证据)**:
|
||||
- 三路:**等权基线**(配置级零开发)/**ICIR 在位者**(quant12_v2a)/**LightGBM challenger**(新开)——同池同窗同考核,#69 影子 A/B 既有基建+归因 challenger/ 子目录机制复用;
|
||||
- LGBM challenger 要素:qlib Alpha158 基准超参起步(lambda_l1=205.7/lambda_l2=581/lr=0.1/num_leaves=210/min_child_samples=100/early_stopping=50);**月度重训**对齐月度批评节奏(优于 QuantaAlpha 一次训练跑四年 test 的弱点);特征=晋级因子池双侧 CSRankNorm 截面秩归一;label=T+1→T+2 防缺口口径;
|
||||
- **池贪心准入(自建;QuantaAlpha 论文有代码无——实验脚本未开源)**:近窗 RankIC 贪心准入+|corr|<0.7 去冗余(高相关对弃低 RankIC)+容量上限(起步=晋级因子全量不设限;饱和点意识挂账——池到几十个再议上限,论文饱和点≈350 因子参考);
|
||||
- 判读:等权≈ICIR=因子相关性高、加权信息少,安心用简单;LGBM 明显跑赢两者=组合方式升级叙事,且 feature importance 天然产出「组合内各因子贡献」(对接 §4.7 归因);预期两三个完整月度周期才有判读力;
|
||||
- 纪律:challenger 永远影子不进生产;若升级仍走决议 K 考场晋级+用户口令,与实盘升级同构。
|
||||
- **班次二落地(2026-10-10)**:`sanguo_factor/challenger_lgbm.py`——v1 池=quant12 在役 12 源(扩池挂后续);三路=方向调整等权/ICIR 档案加权/LGBM(qlib 基准超参,固定 500 轮——无独立 valid 段不早停,注记在对拍件);walk-forward=前 11M 训练+末 1M 样外(test 隔离铁律:样外月绝不进训练;codex 复审加固=训练窗末 2 交易日 embargo——build_label 双 shift 前视,不 embargo 则窗口末样本的 label 吃样外月收盘=信息层泄漏;三路统一可评样本 mask、缺源/坏件整月 skip 不产半吊子件);对拍件=`challenger_lgbm/{host}_{as_of}.json`(三路样外 IC/ICIR/Q5-Q1 多空+feature importance+池冗余注记——`|corr|>0.7` 只记录不剔除,首年观察期剔除留人);挂月度链 stage6(失败不阻链);判读需累积≥2-3 个样外月。**宽表源现场核实修正**:月度批 `values/` 导出只含注册表因子(fa 族+composite),永不含 quant12 12 源——stage6 `--values-dir` 改指 stage4 截面盘 `/volume1/stock/factor_cross_section`(daily_section 幂等 merge 累积的 12 源主 parquet,12/12 在位实证),零重算同链上游供给。
|
||||
- **test 隔离条款(2026-10-10 成文,QuantaAlpha 附录 E 纪律借鉴)**:**合成层(含 challenger)权重/训练只用截至 as_of 的数据,表现评估只看 as_of 之后——权重窗与成绩窗永不重叠**。现状核实动作=审计 quant12_v2a 权重计算代码确认滚动结构天然合规(每月 T-12~T 算权、T+1 后出成绩),核实结论落本档。防什么=将来实盘与回测打架时,回测成绩有「未被自己污染」的底气;防将来改动(全历史拟合/未来数据调参类)悄悄破坏。
|
||||
核实结论(2026-10-10 审计):quant12_v2a(quant12_icirfit_v1)fit_window=2018-01~2021-06、h2 考场窗=2021-07-01 起(exam_gate DEFAULT_START,同批同 run 对照)——**无重叠**(考场首日恰后于拟合窗末月,首尾相接);v2c 滚动天然合规:Y 年制度档案 fit [Y-3 年初, Y-1 年末]、考核 Y 年段(walk-forward),且注册路径只挂载档案冻结权重(load_profile→resolve_weights),评估数据不回流拟合。
|
||||
- **版本钉住+影子对照闸门(决议 A)**:策略钉住合成层具体版本;合成层变更(晋级/退役/调权)→新旧版影子对照(起步默认 4 周,首年校准)→用户口令→各策略**择期**升级。可回滚、历史可复现,与推 vps 纪律同构。
|
||||
- **集体衰减判别门(决议 C)**:月度批评输出两个指标——个别因子衰减(相对同侪,E 环照常机械告警)+集体水位(全体 IC 中位数)。**集体水位跌破阈→不开降权 issue,改开「市况研判」卡片上交用户裁决,因子权重冻结不动。**集体信号=判断题归人,个别信号=机械题归系统;防机械追涨杀跌。
|
||||
- **研判卡=带介入菜单的决策卡**:七档介入阶梯(从轻到重)——
|
||||
|
||||
@@ -60,6 +60,7 @@ export interface PipelineFactor {
|
||||
promotedAtT: number | null
|
||||
decayMonths: number
|
||||
lastEvalDate: string
|
||||
similarity?: Array<{ name: string; ratio: number }> | null // 疑似换皮提示(注册时落)
|
||||
}
|
||||
|
||||
export interface MonthlyReview {
|
||||
@@ -201,6 +202,9 @@ export interface FactorDetail {
|
||||
createdAt: string | null
|
||||
version: string
|
||||
icRecentT: number | null
|
||||
icStats?: { icAll: number | null; tAll: number | null;
|
||||
positiveRatio: number | null; byYear: Record<string, number> } | null
|
||||
similarity?: Array<{ name: string; ratio: number }> | null
|
||||
lastEvalDate: string
|
||||
}
|
||||
birth: { jobId: string; startedAt: string; rounds: number | null; hypId: string } | null
|
||||
|
||||
@@ -131,3 +131,42 @@ describe('FactorDetail.vue 危险区送墓园(2026-10-09 补丁)', () => {
|
||||
'insider_buy_decay60', 'graveyard', [], '衰减不止')
|
||||
})
|
||||
})
|
||||
|
||||
describe('FactorDetail.vue 全期弹药+疑似换皮(spec §4.2 2026-10-10)', () => {
|
||||
beforeEach(() => vi.clearAllMocks())
|
||||
it('渲染弹药行(全期IC/合成t/正IC占比/逐年)与换皮 chips', async () => {
|
||||
const w = await mountDetail({
|
||||
factor: { ...DETAIL.factor,
|
||||
icStats: { icAll: 0.04, tAll: 1.73, positiveRatio: 1.0,
|
||||
byYear: { 2025: 0.05, 2026: 0.02 } },
|
||||
similarity: [{ name: 'fa_old', ratio: 1.0 }] },
|
||||
birth: null,
|
||||
})
|
||||
expect(w.text()).toContain('全期IC 0.0400')
|
||||
expect(w.text()).toContain('合成t 1.73')
|
||||
expect(w.text()).toContain('正IC占比 100%')
|
||||
expect(w.text()).toContain('2025:0.0500')
|
||||
expect(w.text()).toContain('2026:0.0200')
|
||||
expect(w.findAll('.sim-chip').length).toBe(1)
|
||||
expect(w.text()).toContain('fa_old')
|
||||
})
|
||||
it('无 icStats(月度点不足或首班未跑)显示 — 占位,不造数', async () => {
|
||||
const w = await mountDetail()
|
||||
expect(w.text()).toContain('月度点不足或首班未跑')
|
||||
expect(w.findAll('.sim-chip').length).toBe(0)
|
||||
})
|
||||
it('codex review:tAll=null(样本不足/方差0)时其余弹药字段照常渲染,合成t 单独 —', async () => {
|
||||
const w = await mountDetail({
|
||||
factor: { ...DETAIL.factor,
|
||||
icStats: { icAll: 0.04, tAll: null, positiveRatio: 0.8,
|
||||
byYear: { 2025: 0.05 } } },
|
||||
birth: null,
|
||||
})
|
||||
expect(w.text()).toContain('全期IC 0.0400')
|
||||
expect(w.text()).toContain('合成t —')
|
||||
expect(w.text()).toContain('正IC占比 80%')
|
||||
expect(w.text()).toContain('2025:0.0500')
|
||||
// 区分「无报告」:此态不再出占位文案
|
||||
expect(w.text()).not.toContain('月度点不足或首班未跑')
|
||||
})
|
||||
})
|
||||
|
||||
@@ -67,6 +67,11 @@ async function sendToGraveyard(): Promise<void> {
|
||||
ElMessage.error(`判死失败: ${e instanceof Error ? e.message : e}`)
|
||||
} finally { burying.value = false }
|
||||
}
|
||||
|
||||
// 疑似换皮 chip 点击跳对方详情(比对靠人,一键可达)
|
||||
function goFactor(n: string): void {
|
||||
router.push(`/pipeline/factors/${n}`)
|
||||
}
|
||||
</script>
|
||||
|
||||
<template>
|
||||
@@ -119,6 +124,22 @@ async function sendToGraveyard(): Promise<void> {
|
||||
<div v-if="isLegacy" class="brow"><span class="bk">档案说明</span>
|
||||
<span class="dim">此因子出生于档案制度(2026-10-09)之前,表达式/自述/承载逻辑未随出生落盘——存量形态,非数据丢失</span>
|
||||
</div>
|
||||
<div class="brow"><span class="bk">全期弹药</span>
|
||||
<span class="val">
|
||||
<template v-if="detail.factor.icStats">
|
||||
全期IC {{ detail.factor.icStats.icAll?.toFixed(4) ?? '—' }} ·
|
||||
合成t <span :title="detail.factor.icStats.tAll === null ? 'n<2 或方差为 0,如实不算' : ''">{{ detail.factor.icStats.tAll !== null && detail.factor.icStats.tAll !== undefined ? detail.factor.icStats.tAll.toFixed(2) : '—' }}</span> ·
|
||||
正IC占比 {{ detail.factor.icStats.positiveRatio !== null && detail.factor.icStats.positiveRatio !== undefined ? `${(detail.factor.icStats.positiveRatio * 100).toFixed(0)}%` : '—' }} ·
|
||||
逐年 {{ Object.entries(detail.factor.icStats.byYear).map(([y, v]) => `${y}:${v.toFixed(4)}`).join(' ') || '—' }}
|
||||
</template>
|
||||
<template v-else>—(月度点不足或首班未跑)</template>
|
||||
</span></div>
|
||||
<div class="brow" v-if="detail.factor.similarity?.length"><span class="bk">疑似换皮</span>
|
||||
<span class="val warn">
|
||||
<span v-for="s in detail.factor.similarity" :key="s.name"
|
||||
class="sim-chip" @click="goFactor(s.name)">
|
||||
{{ s.name }} · {{ (s.ratio * 100).toFixed(0) }}%</span>
|
||||
</span></div>
|
||||
</PanelCard>
|
||||
|
||||
<div class="stats-row">
|
||||
@@ -169,4 +190,8 @@ async function sendToGraveyard(): Promise<void> {
|
||||
.empty-box { color: var(--text-3); font-size: 12px; padding: 14px 4px; }
|
||||
/* 危险区:送墓园 */
|
||||
.grave-row { display: flex; align-items: center; gap: 10px; flex-wrap: wrap; }
|
||||
/* 疑似换皮 chips */
|
||||
.val.warn { color: var(--warn, #e6a23c); }
|
||||
.sim-chip { cursor: pointer; border: 1px solid rgba(230, 162, 60, 0.45); background: rgba(230, 162, 60, 0.08); border-radius: 5px; padding: 0 6px; margin-right: 6px; line-height: 18px; display: inline-block; font-size: 11px; }
|
||||
.sim-chip:hover { border-color: var(--warn, #e6a23c); }
|
||||
</style>
|
||||
|
||||
@@ -69,3 +69,24 @@ describe('FactorFactory.vue', () => {
|
||||
expect(w.findAll('.src-none').length).toBe(factorsMock.filter((f) => !f.hypothesis).length)
|
||||
})
|
||||
})
|
||||
|
||||
describe('FactorFactory.vue 疑似换皮徽标(spec §4.2 2026-10-10)', () => {
|
||||
beforeEach(() => {
|
||||
vi.clearAllMocks()
|
||||
vi.mocked(api.getFactors).mockResolvedValue(factorsMock)
|
||||
vi.mocked(api.getMonthlyReview).mockResolvedValue(monthlyReviewMock)
|
||||
})
|
||||
it('similarity 落款的行出 ⚠ 徽标,无落款不出', async () => {
|
||||
const rows = factorsMock.map((f, i) => i === 0
|
||||
? { ...f, similarity: [{ name: 'fa_old', ratio: 0.9 }] } : f)
|
||||
vi.mocked(api.getFactors).mockResolvedValueOnce(rows)
|
||||
const w = mount(FactorFactory, { global: { plugins: [ElementPlus] } })
|
||||
await flushPromises()
|
||||
expect(w.findAll('.dup-badge').length).toBe(1)
|
||||
})
|
||||
it('全部无 similarity 时零徽标', async () => {
|
||||
const w = mount(FactorFactory, { global: { plugins: [ElementPlus] } })
|
||||
await flushPromises()
|
||||
expect(w.findAll('.dup-badge').length).toBe(0)
|
||||
})
|
||||
})
|
||||
|
||||
@@ -69,6 +69,13 @@ const fmt = (v: number | null): string => (v === null ? '—' : v.toFixed(2))
|
||||
<a class="src-link" @click.stop="router.push('/pipeline/hypotheses')">🔗</a>
|
||||
</el-tooltip>
|
||||
<span v-else class="src-none">—</span>
|
||||
<!-- AST 防换皮徽标(spec §4.2 2026-10-10):注册时相似度≥0.6 落
|
||||
similarity,列表侧只提示——点行进详情看比对 chips,判定人做 -->
|
||||
<el-tooltip v-if="row.similarity?.length"
|
||||
:content="`疑似换皮: ${row.similarity.map((s: { name: string }) => s.name).join(', ')}`"
|
||||
placement="top">
|
||||
<span class="dup-badge" @click.stop="router.push(`/pipeline/factors/${row.id}`)">⚠</span>
|
||||
</el-tooltip>
|
||||
</div>
|
||||
</template>
|
||||
</el-table-column>
|
||||
@@ -128,6 +135,8 @@ const fmt = (v: number | null): string => (v === null ? '—' : v.toFixed(2))
|
||||
.src-link { font-size: 10.5px; color: var(--cyan-dim); cursor: pointer; border-bottom: 1px dashed rgba(0, 184, 204, 0.4); }
|
||||
.src-link:hover { color: var(--brand); }
|
||||
.src-none { color: var(--text-3); font-size: 10.5px; }
|
||||
/* 疑似换皮徽标:提示非硬拒,点行进详情看比对 */
|
||||
.dup-badge { cursor: pointer; font-size: 11px; color: var(--warn, #e6a23c); }
|
||||
.el-table { cursor: pointer; }
|
||||
.alert-txt { color: #ff5f6d; font-family: var(--mono); font-size: 11.5px; }
|
||||
.dim { color: var(--text-3); }
|
||||
|
||||
@@ -437,6 +437,13 @@ def factors() -> dict:
|
||||
reg = vr.load_registry(ensure_runtime_registry(
|
||||
_factor_registry_path(), os.path.join("config", "factor_registry.yaml")))
|
||||
points, last_eval = _eval_latest_points()
|
||||
# 全史 t(=全期合成 t)从最新月报 factor_stats 一次性接真(spec §4.2 承诺
|
||||
# factors 端点升级;单次读取,勿循环内重读——codex review HIGH)
|
||||
reports = _monthly_reports()
|
||||
full_t: dict[str, Any] = (
|
||||
{n: (s or {}).get("tAll")
|
||||
for n, s in (reports[0][2].get("factor_stats") or {}).items()}
|
||||
if reports else {})
|
||||
items: list[dict[str, Any]] = []
|
||||
for name, e in reg["factors"].items():
|
||||
if e.get("status") == "graveyard":
|
||||
@@ -453,8 +460,9 @@ def factors() -> dict:
|
||||
"hypothesis": e.get("hypothesis"),
|
||||
"description": e.get("description") or "",
|
||||
"icRecentT": t, # 最近批(12M 滚动窗)即近窗 t
|
||||
"icFullT": None, # eval_store 每批单窗 t,无全史批数据源,如实 null
|
||||
"icFullT": full_t.get(name), # 全期合成 t(月报 factor_stats.tAll)
|
||||
"promotedAtT": e.get("promotion_t") or None,
|
||||
"similarity": e.get("similarity"), # 疑似换皮提示(注册时落)
|
||||
"decayMonths": 0,
|
||||
"lastEvalDate": last_eval})
|
||||
order = {"decaying": 0, "promoted": 1, "assessable": 2,
|
||||
@@ -605,6 +613,10 @@ def factor_detail(name: str) -> dict:
|
||||
birth = {"jobId": job["jobId"], "startedAt": job["startedAt"],
|
||||
"rounds": job["rounds"], "hypId": hyp_id}
|
||||
break
|
||||
stats = None
|
||||
reports = _monthly_reports()
|
||||
if reports:
|
||||
stats = (reports[0][2].get("factor_stats") or {}).get(name)
|
||||
return {"factor": {
|
||||
"name": name, "status": e["status"],
|
||||
"origin": e.get("origin") or "manual", "hypothesis": hyp_id,
|
||||
@@ -615,6 +627,8 @@ def factor_detail(name: str) -> dict:
|
||||
"createdAt": first.get("effective_from"),
|
||||
"version": str(first.get("v")) if first else "1",
|
||||
"icRecentT": (points.get(name) or {}).get("t"),
|
||||
"icStats": stats,
|
||||
"similarity": e.get("similarity"),
|
||||
"lastEvalDate": last_eval}, "birth": birth}
|
||||
|
||||
|
||||
@@ -1162,6 +1176,9 @@ async def _decompose_worker(hyp_id: str, config, job_id: str | None = None) -> d
|
||||
today = date.today().isoformat()
|
||||
for cand in result["passed"]:
|
||||
src = derive_source(cand["expression"]) # 已过门,必为 str
|
||||
# 仅建条纪律(codex review):已存在条目(同名同假设幂等重注册)
|
||||
# 不重写 similarity——与 description/origin 落盘纪律同款
|
||||
is_new = cand["name"] not in reg["factors"]
|
||||
vr.upsert_factor(reg, cand["name"], hypothesis=hyp_id,
|
||||
status="incubating", origin="decomposer",
|
||||
description=cand.get("description"))
|
||||
@@ -1170,6 +1187,18 @@ async def _decompose_worker(hyp_id: str, config, job_id: str | None = None) -> d
|
||||
"source": src, "origin": "decomposer",
|
||||
"justification": cand["justification"]},
|
||||
effective_from=today)
|
||||
# AST 防换皮提示(spec §4.2 2026-10-10):对既有池比对,≥0.6 存
|
||||
# entry.similarity——硬门(dup_subtree)外的部分同构仍可注册,
|
||||
# 提示非硬拒,判定人做.池=注册表条目当前表达式(versions[-1]).
|
||||
if is_new:
|
||||
from sanguo_factor import expression_match as em
|
||||
_cands = {n: str(((v.get("versions") or [{}])[-1]
|
||||
.get("params") or {}).get("expression") or "")
|
||||
for n, v in reg["factors"].items()
|
||||
if n != cand["name"]}
|
||||
_sim = em.top_similar(str(cand["expression"]), _cands)
|
||||
if _sim:
|
||||
reg["factors"][cand["name"]]["similarity"] = _sim
|
||||
registered.append({"name": cand["name"], "source": src,
|
||||
"expression": cand["expression"],
|
||||
"description": cand.get("description", ""),
|
||||
|
||||
@@ -0,0 +1,323 @@
|
||||
# sanguo_factor/challenger_lgbm.py
|
||||
"""LGBM challenger 三路加权对拍(2026-10-10 spec §4.4 定案).
|
||||
|
||||
三路=等权(方向调整)/ICIR(quant12_icirfit_v1 档案)/LightGBM(月度重训),
|
||||
同池同窗 walk-forward 影子对拍;challenger 永不进生产,对拍件落
|
||||
reports/factor_monthly/challenger_lgbm/ 子目录(判定层端点只读根层).
|
||||
对拍件同 host+as_of 幂等覆盖(月度重跑覆写同 key).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
import pandas as pd
|
||||
|
||||
_PROFILE = Path(__file__).parent / "weight_profiles" / "quant12_icirfit_v1.json"
|
||||
|
||||
|
||||
def load_pool() -> dict[str, dict]:
|
||||
"""v1 对拍池=在役 12 源(ICIR 档案现成,三路同池才可比;扩池挂后续)."""
|
||||
with open(_PROFILE, encoding="utf-8") as f:
|
||||
return json.load(f)["sources"]
|
||||
|
||||
|
||||
def build_label(close_wide: pd.DataFrame) -> pd.DataFrame:
|
||||
"""防缺口 label=T+1 收盘→T+2 收盘(信号 T 收盘出,T+1 全天可成交)."""
|
||||
return close_wide.shift(-2) / close_wide.shift(-1) - 1.0
|
||||
|
||||
|
||||
def cs_rank_norm(df: pd.DataFrame) -> pd.DataFrame:
|
||||
"""截面秩归一 (rank_pct-0.5);NaN 透传不占当日截面."""
|
||||
return df.rank(axis=1, pct=True) - 0.5
|
||||
|
||||
|
||||
LGBM_PARAMS = { # qlib Alpha158 基准超参起步(spec §4.4;无独立 valid 段,v1 固定轮数)
|
||||
"objective": "mse", "learning_rate": 0.1, "max_depth": 8, "num_leaves": 210,
|
||||
"colsample_bytree": 0.8879, "subsample": 0.8789,
|
||||
"lambda_l1": 205.6999, "lambda_l2": 580.9768,
|
||||
"min_child_samples": 100, "feature_fraction_bynode": 0.8,
|
||||
"seed": 42, "num_threads": 4, "verbose": -1,
|
||||
}
|
||||
NUM_BOOST_ROUND = 500
|
||||
|
||||
|
||||
def _aligned(values: dict[str, pd.DataFrame]) -> pd.DataFrame:
|
||||
"""多因子宽表纵向拼接成特征长表(index 对齐,缺失源列 NaN)."""
|
||||
return pd.concat({n: cs_rank_norm(v) for n, v in values.items()}, axis=1)
|
||||
|
||||
|
||||
def _dir_sign(pool: dict, name: str) -> float:
|
||||
return -1.0 if pool.get(name, {}).get("direction") == "-" else 1.0
|
||||
|
||||
|
||||
def equal_weight_signal(values: dict[str, pd.DataFrame], pool: dict) -> pd.DataFrame:
|
||||
stack = pd.concat([_dir_sign(pool, n) * cs_rank_norm(v) for n, v in values.items()])
|
||||
return stack.groupby(level=0).mean()
|
||||
|
||||
|
||||
def icir_signal(values: dict[str, pd.DataFrame], pool: dict) -> pd.DataFrame:
|
||||
total = sum(pool[n].get("weight", 0.0) for n in values)
|
||||
if total <= 0:
|
||||
raise ValueError("ICIR 权重和为零,档案异常")
|
||||
out = None
|
||||
for n, v in values.items():
|
||||
w = pool[n].get("weight", 0.0) * _dir_sign(pool, n)
|
||||
part = w * cs_rank_norm(v)
|
||||
out = part if out is None else out.add(part, fill_value=0.0)
|
||||
return out / total
|
||||
|
||||
|
||||
def _split_masks(index: pd.DatetimeIndex, last_month: str,
|
||||
embargo_days: int = 2, train_months: int = 11) -> tuple[np.ndarray, np.ndarray]:
|
||||
"""walk-forward 切分 mask(codex 复审 H-1/H→M-7).
|
||||
|
||||
oos=last_month 当月;train=[last_month 前推 train_months 个自然月之首,
|
||||
oos 首日前 embargo_days 个交易日)。embargo 保证:训练样本 t 的 label
|
||||
(close[t+2]/close[t+1]) 取值严格早于样外月首日——label 信息层不泄漏
|
||||
(build_label 双 shift 前视,窗口末 2 交易的 label 会吃到样外月收盘).
|
||||
训练下界显式裁剪对齐 spec 12M 语义,防截面盘逐年累积导致训练窗漂移.
|
||||
"""
|
||||
oos = np.asarray(index.strftime("%Y-%m")) == last_month
|
||||
train = np.zeros(len(index), dtype=bool)
|
||||
if not oos.any():
|
||||
return train, oos
|
||||
first_oos = int(np.argmax(oos))
|
||||
train[:max(first_oos - embargo_days, 0)] = True
|
||||
floor = (pd.Period(last_month, freq="M") - train_months).to_timestamp()
|
||||
train &= np.asarray(index >= floor)
|
||||
return train, oos
|
||||
|
||||
|
||||
def lgbm_walk_forward(values: dict[str, pd.DataFrame], label: pd.DataFrame,
|
||||
last_month: str, params: dict | None = None) -> tuple[pd.DataFrame, dict]:
|
||||
"""训练窗=样外月前推 11 个自然月,末 2 交易日 embargo;样外=last_month 当月.
|
||||
|
||||
返回 (样外日×股票预测宽表, {feature: gain}).特征=CSRankNorm 后各源,
|
||||
label=CSRankNorm 后防缺口收益;日频截面样本(日期,股票)平铺训练.
|
||||
"""
|
||||
import lightgbm as lgb
|
||||
|
||||
feat = _aligned(values)
|
||||
lab = cs_rank_norm(label)
|
||||
common = feat.index.intersection(lab.index)
|
||||
feat, lab = feat.loc[common], lab.loc[common]
|
||||
|
||||
train_mask, oos_mask = _split_masks(feat.index, last_month)
|
||||
if not oos_mask.any():
|
||||
return pd.DataFrame(), {}
|
||||
X_tr = feat[train_mask].stack(future_stack=True).reset_index()
|
||||
X_tr.columns = ["datetime", "vt_symbol", *feat.columns.levels[0]]
|
||||
y_df = lab.stack(future_stack=True).rename("y").reset_index()
|
||||
y_df.columns = ["datetime", "vt_symbol", "y"] # index 无名时 reset 生成 level_0/1,显式定名
|
||||
tr = X_tr.merge(y_df, on=["datetime", "vt_symbol"]).dropna()
|
||||
if tr.empty:
|
||||
return pd.DataFrame(), {}
|
||||
X = tr[list(feat.columns.levels[0])]
|
||||
model = lgb.train(params or LGBM_PARAMS, lgb.Dataset(X, label=tr["y"]),
|
||||
num_boost_round=NUM_BOOST_ROUND)
|
||||
imp = dict(zip(X.columns, model.feature_importance("gain").tolist()))
|
||||
|
||||
X_oos = feat[oos_mask].stack(future_stack=True).reset_index()
|
||||
X_oos.columns = X_tr.columns
|
||||
X_oos = X_oos.dropna(subset=list(feat.columns.levels[0]))
|
||||
X_oos = X_oos.merge(y_df.dropna()[["datetime", "vt_symbol"]],
|
||||
on=["datetime", "vt_symbol"]) # 只留 label 有效对(可评分样外)
|
||||
if X_oos.empty:
|
||||
return pd.DataFrame(), imp
|
||||
preds = model.predict(X_oos[list(feat.columns.levels[0])])
|
||||
out = X_oos[["datetime", "vt_symbol"]].assign(p=preds)
|
||||
return out.pivot(index="datetime", columns="vt_symbol", values="p"), imp
|
||||
|
||||
|
||||
def _load_close_wide(vnpy_db: str, columns_ref: pd.Index, start: str, end: str) -> pd.DataFrame:
|
||||
"""按因子宽表列(股票)拉收盘价(universe.load_universe_bars 轻量列).
|
||||
|
||||
真实 loader 必传 start/end(universe.py);+45d 前向缓冲覆盖 label 的 t+2。
|
||||
直读前 purge 同窗缓存(daily_section 同款纪律:防旧缓存把新到 bar 截在
|
||||
缓存生成日)——challenger 月度重跑不复用旧缓存件,选型=每次全量重读,
|
||||
月频成本可接受。
|
||||
"""
|
||||
from .universe import load_universe_bars, purge_cache
|
||||
|
||||
db = vnpy_db
|
||||
if not db:
|
||||
from sanguo_data.config import find_config_path, load_config
|
||||
db = load_config(find_config_path()).data_paths["vnpy_db"]
|
||||
purge_cache(db, start, end)
|
||||
bars = load_universe_bars(db, start, end, symbols=list(columns_ref))
|
||||
_ = bars.select(["datetime", "vt_symbol", "close"]).to_pandas()
|
||||
wide = _.pivot(index="datetime", columns="vt_symbol", values="close").sort_index()
|
||||
wide.index = pd.to_datetime(wide.index)
|
||||
return wide
|
||||
|
||||
|
||||
def score_signal(signal: pd.DataFrame, label: pd.DataFrame) -> dict:
|
||||
"""样外评分:日 IC 均值/ICIR/五分位多空累计/有效天数.
|
||||
|
||||
days=进入评分的天数(截面≥5 对)——非有效 IC 天数(常数信号 IC=NaN 也计入).
|
||||
q5q1=20%/80% 分位阈值的近似多空累计(tie 语义:quantile 阈值含等值边界,
|
||||
小截面下是近似而非严格五分位分组);IC 全 NaN 时 ic_mean/icir=None.
|
||||
"""
|
||||
common = signal.index.intersection(label.index)
|
||||
ics = []
|
||||
ls_rets = []
|
||||
for d in common:
|
||||
s, y = signal.loc[d], label.loc[d]
|
||||
pair = pd.concat([s, y], axis=1, keys=["s", "y"]).dropna()
|
||||
if len(pair) < 5:
|
||||
continue
|
||||
ics.append(pair["s"].corr(pair["y"], method="spearman"))
|
||||
q = pair["s"].quantile([0.2, 0.8])
|
||||
lo, hi = pair[pair["s"] <= q[0.2]]["y"].mean(), pair[pair["s"] >= q[0.8]]["y"].mean()
|
||||
ls_rets.append((hi - lo) if (lo is not None and hi is not None) else 0.0)
|
||||
if not ics:
|
||||
return {"ic_mean": None, "icir": None, "q5q1": None, "days": 0}
|
||||
ser = pd.Series(ics).dropna()
|
||||
icir = (ser.mean() / ser.std()) if len(ser) > 1 and ser.std() > 0 else None
|
||||
return {"ic_mean": round(float(ser.mean()), 6) if len(ser) else None,
|
||||
"icir": round(float(icir), 6) if icir is not None else None,
|
||||
"q5q1": round(float(sum(ls_rets)), 6), "days": len(ics)}
|
||||
|
||||
|
||||
def _read_value_parquet(path: str) -> pd.DataFrame | None:
|
||||
"""读因子宽表 parquet;坏件/空件/索引不可转 datetime → None(codex 复审 M-5).
|
||||
|
||||
结构化剔除走缺源 skip 语义,不裸抛 AttributeError/ArrowInvalid.
|
||||
索引须严格升序且唯一(codex 尾批):embargo 的 shift(-2) 以位置序为地基,
|
||||
乱序/重复日期件拒收.
|
||||
"""
|
||||
try:
|
||||
df = pd.read_parquet(path)
|
||||
if df.empty or df.columns.empty:
|
||||
return None
|
||||
idx = pd.to_datetime(df.index)
|
||||
if not (idx.is_monotonic_increasing and idx.is_unique):
|
||||
print(f"[challenger] ⚠️ parquet 索引乱序/重复: {path} "
|
||||
f"(monotonic={idx.is_monotonic_increasing}, unique={idx.is_unique})")
|
||||
return None
|
||||
except Exception as exc:
|
||||
print(f"[challenger] ⚠️ parquet 不可读: {path} ({exc!r})")
|
||||
return None
|
||||
return df.set_axis(idx)
|
||||
|
||||
|
||||
def _redundancy_flags(values: dict[str, pd.DataFrame], threshold: float = 0.7) -> list[dict]:
|
||||
"""池冗余注记:按日源间 corr 均值 |corr|>threshold 只记录不剔除(首年观察期).
|
||||
|
||||
按日循环构造 股票×源 帧后 corr=源×源(codex 复审 M-8:替代全历史 stack 的
|
||||
MultiIndex 大表——12 源×12M×全A 在 NAS 2G 内存机不可行);逐对 NaN 感知
|
||||
累积(缺数据日只计入有值对).
|
||||
"""
|
||||
names = list(values)
|
||||
normed = {n: cs_rank_norm(v) for n, v in values.items()}
|
||||
all_dates = sorted(set().union(*(set(v.index) for v in normed.values())))
|
||||
sum_c = None
|
||||
cnt_c = None
|
||||
for d in all_dates:
|
||||
day = pd.DataFrame({n: normed[n].loc[d] for n in names if d in normed[n].index})
|
||||
if day.shape[0] < 2 or day.shape[1] < 2:
|
||||
continue
|
||||
c = day.corr(method="spearman")
|
||||
sum_c = c if sum_c is None else sum_c.add(c, fill_value=0.0)
|
||||
cnt = c.notna().astype(float)
|
||||
cnt_c = cnt if cnt_c is None else cnt_c.add(cnt, fill_value=0.0)
|
||||
flagged = []
|
||||
if sum_c is None:
|
||||
return flagged
|
||||
means = sum_c / cnt_c.replace(0.0, np.nan)
|
||||
for n1 in names:
|
||||
for n2 in names:
|
||||
if n1 < n2 and n1 in means.index and n2 in means.columns:
|
||||
v = means.loc[n1, n2]
|
||||
if pd.notna(v) and abs(v) > threshold:
|
||||
flagged.append({"a": n1, "b": n2, "corr": round(float(v), 4)})
|
||||
return flagged
|
||||
|
||||
|
||||
def run_challenge(values_dir: str, vnpy_db: str, as_of: str, out_dir: str,
|
||||
host: str = "nas") -> str:
|
||||
"""月度链 stage6 入口:读导出宽表→三路→对拍件(同 host+as_of 幂等覆盖).
|
||||
|
||||
池完整性铁律(codex 复审 H-3):任何源缺件/坏件 → FileNotFoundError 走
|
||||
main 的 skip 路径,不产半吊子对拍件.
|
||||
"""
|
||||
pool = load_pool()
|
||||
names = sorted(pool)
|
||||
loaded: dict[str, pd.DataFrame] = {}
|
||||
for n in names:
|
||||
p = os.path.join(values_dir, f"{n}.parquet")
|
||||
if not os.path.exists(p):
|
||||
continue
|
||||
df = _read_value_parquet(p)
|
||||
if df is not None:
|
||||
loaded[n] = df
|
||||
missing = [n for n in names if n not in loaded]
|
||||
if missing:
|
||||
raise FileNotFoundError(
|
||||
f"missing sources: expected={len(pool)} loaded={len(loaded)} "
|
||||
f"missing={missing} (values_dir={values_dir})")
|
||||
|
||||
cols_ref = pd.Index(sorted(set().union(*(set(v.columns) for v in loaded.values()))))
|
||||
start = min(v.index.min() for v in loaded.values()).strftime("%Y-%m-%d")
|
||||
end = max(v.index.max() for v in loaded.values()).strftime("%Y-%m-%d")
|
||||
close = _load_close_wide(vnpy_db, cols_ref, start, end)
|
||||
label = build_label(close)
|
||||
last_month = as_of[:7]
|
||||
|
||||
signals = {"equal": equal_weight_signal(loaded, pool),
|
||||
"icir": icir_signal(loaded, pool)}
|
||||
pred, imp = lgbm_walk_forward(loaded, label, last_month)
|
||||
if not pred.empty:
|
||||
signals["lgbm"] = pred
|
||||
|
||||
# 统一可评样本 mask(codex 复审 H-2):OOS 月内全部源非空且 label 非空的
|
||||
# date×stock 集合=唯一评分集,三路全部 reindex 到该集合再 score
|
||||
oos_label = label[label.index.strftime("%Y-%m") == last_month]
|
||||
mask = pd.DataFrame(True, index=oos_label.index, columns=cols_ref)
|
||||
for v in loaded.values():
|
||||
mask &= v.reindex(index=oos_label.index, columns=cols_ref).notna()
|
||||
mask &= oos_label.reindex(index=oos_label.index, columns=cols_ref).notna()
|
||||
scored = {k: score_signal(sig.reindex(index=oos_label.index, columns=cols_ref).where(mask),
|
||||
oos_label)
|
||||
for k, sig in signals.items()}
|
||||
|
||||
flagged = _redundancy_flags(loaded)
|
||||
|
||||
doc = {"as_of": as_of, "generated_at": pd.Timestamp.now().isoformat(),
|
||||
"host": host, "pool": {"names": sorted(loaded), "size": len(loaded)},
|
||||
"split": {"oos_month": last_month,
|
||||
"note": "训练=样外月前推 11 个自然月(末 2 交易日 embargo,"
|
||||
"label 信息层隔离);无独立 valid,固定轮数"},
|
||||
"signals": scored, "feature_importance": imp,
|
||||
"redundancy_flagged": flagged,
|
||||
"lgbm_params": {**LGBM_PARAMS, "num_boost_round": NUM_BOOST_ROUND}}
|
||||
sub = os.path.join(out_dir, "challenger_lgbm")
|
||||
os.makedirs(sub, exist_ok=True)
|
||||
path = os.path.join(sub, f"{host}_{as_of}.json")
|
||||
with open(path, "w", encoding="utf-8") as f:
|
||||
json.dump(doc, f, ensure_ascii=False, indent=2, allow_nan=False)
|
||||
return path
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
import argparse
|
||||
ap = argparse.ArgumentParser(description="LGBM challenger 三路对拍(月度链 stage6)")
|
||||
ap.add_argument("--values-dir", required=True,
|
||||
help="因子截面宽表目录(每因子一 parquet;NAS=stage4 截面盘"
|
||||
" /volume1/stock/factor_cross_section,月度批 values/ 只含"
|
||||
"注册表因子不含 quant12 12 源)")
|
||||
ap.add_argument("--as-of", required=True)
|
||||
ap.add_argument("--out-dir", required=True, help="报告根(factor_monthly)")
|
||||
ap.add_argument("--vnpy-db", default=None)
|
||||
ap.add_argument("--host", default="nas")
|
||||
a = ap.parse_args(argv)
|
||||
try:
|
||||
path = run_challenge(a.values_dir, a.vnpy_db, a.as_of, a.out_dir, host=a.host)
|
||||
except FileNotFoundError as e:
|
||||
print(f"[challenger] skip: {e}")
|
||||
return 0 # 无宽表=非错误(截面盘未建),不阻链
|
||||
print(f"[challenger] 对拍件: {path}")
|
||||
return 0
|
||||
@@ -0,0 +1,120 @@
|
||||
# sanguo_factor/expression_match.py
|
||||
"""表达式结构相似度(AST 防换皮,2026-10-10 spec §4.2).
|
||||
|
||||
QuantaAlpha factor_ast 算法思想(最大公共子树+交换律)的方言内实现:
|
||||
不移植其 qlib 式解析器,直接用 Python ast——与 factor_guard 白名单
|
||||
同方言,注册的新因子表达式必然可解析.判定永远人做:本模块只产提示.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import ast
|
||||
|
||||
_COMMUTATIVE = (ast.Add, ast.Mult)
|
||||
|
||||
|
||||
def _is_meta(node: ast.AST) -> bool:
|
||||
"""元数据子节点:expr_context(Load/Store)与算子(Div/Mult/USub...)——
|
||||
前者是 Name 的语境标记,后者的类型已编码进 BinOp/UnaryOp 的 key.
|
||||
不跳过则 ? 兜底签名全局撞车+虚增计数(Name 算 2/整树多 1)."""
|
||||
return isinstance(node, (ast.expr_context, ast.operator, ast.unaryop))
|
||||
|
||||
|
||||
def _node_size(node: ast.AST) -> int:
|
||||
return 1 + sum(_node_size(c) for c in ast.iter_child_nodes(node)
|
||||
if not _is_meta(c))
|
||||
|
||||
|
||||
def _key(node: ast.AST) -> str:
|
||||
"""规范化结构签名:交换律 binop 左右子树排序后拼接.
|
||||
|
||||
兜底分支递归编码全部非元数据子节点——同类型未知节点只有子树也
|
||||
同构才撞签,不再 `?TypeName` 全局等价(codex review CRITICAL).
|
||||
"""
|
||||
if isinstance(node, ast.BinOp) and isinstance(node.op, _COMMUTATIVE):
|
||||
lk, rk = _key(node.left), _key(node.right)
|
||||
a, b = sorted((lk, rk))
|
||||
return f"({a}|{b}|{type(node.op).__name__})"
|
||||
if isinstance(node, ast.BinOp):
|
||||
return f"({_key(node.left)}>{_key(node.right)}|{type(node.op).__name__})"
|
||||
if isinstance(node, ast.UnaryOp):
|
||||
return f"{type(node.op).__name__}({_key(node.operand)})"
|
||||
if isinstance(node, ast.Call) and isinstance(node.func, ast.Name):
|
||||
args = ",".join(_key(a) for a in node.args)
|
||||
kws = ",".join(f"{k.arg}={_key(k.value)}"
|
||||
for k in sorted(node.keywords, key=lambda k: k.arg or ""))
|
||||
return f"{node.func.id}({args};{kws})"
|
||||
if isinstance(node, ast.Name):
|
||||
return f"#{node.id}"
|
||||
if isinstance(node, ast.Constant):
|
||||
return f"#{node.value!r}"
|
||||
kids = ",".join(_key(c) for c in ast.iter_child_nodes(node)
|
||||
if not _is_meta(c))
|
||||
return f"?{type(node).__name__}<{kids}>"
|
||||
|
||||
|
||||
def _call_func_ids(node: ast.AST) -> set[int]:
|
||||
"""Call 的 func Name 节点 id 集:算子名已编码进 Call 签名,不再作为
|
||||
独立子树参与匹配——否则仅共享算子名(ts_mean/f/cs_rank)也计公共
|
||||
子树,产生 1/N 的地板噪声(codex review 修批实测)."""
|
||||
return {id(c.func) for c in ast.walk(node) if isinstance(c, ast.Call)}
|
||||
|
||||
|
||||
def _all_subtree_keys(node: ast.AST) -> dict[str, int]:
|
||||
"""子树签名→节点数(同签名取最大)."""
|
||||
skip = _call_func_ids(node)
|
||||
out: dict[str, int] = {}
|
||||
for sub in ast.walk(node):
|
||||
if _is_meta(sub) or id(sub) in skip:
|
||||
continue
|
||||
k = _key(sub)
|
||||
out[k] = max(out.get(k, 0), _node_size(sub))
|
||||
return out
|
||||
|
||||
|
||||
def similarity(expr_a: str, expr_b: str) -> float | None:
|
||||
"""最大公共子树节点数 / 较小表达式节点数;解析失败或深树爆栈=None.
|
||||
|
||||
取 .body 剥掉 Expression 包装节点——否则其签名落 ? 兜底桶,
|
||||
任意两表达式的根都会撞签(子树=全树,相似度恒 1.0).
|
||||
RecursionError 防护:factor_guard _DEPTH_MAX=12 挡真实注册深树,
|
||||
此处捕爆栈如实 None,不 raise(codex review LOW).
|
||||
"""
|
||||
try:
|
||||
ta = ast.parse(expr_a, mode="eval").body
|
||||
tb = ast.parse(expr_b, mode="eval").body
|
||||
sa, sb = _node_size(ta), _node_size(tb)
|
||||
if sa == 0 or sb == 0:
|
||||
return None
|
||||
keys_a = _all_subtree_keys(ta)
|
||||
skip_b = _call_func_ids(tb)
|
||||
best = 0
|
||||
for sub in ast.walk(tb):
|
||||
if _is_meta(sub) or id(sub) in skip_b:
|
||||
continue
|
||||
k = _key(sub)
|
||||
if k in keys_a:
|
||||
best = max(best, min(keys_a[k], _node_size(sub)))
|
||||
return round(best / min(sa, sb), 4)
|
||||
except (SyntaxError, ValueError, RecursionError):
|
||||
return None
|
||||
|
||||
|
||||
def top_similar(expression: str, candidates: dict[str, str],
|
||||
floor: float = 0.6, limit: int = 5) -> list[dict]:
|
||||
"""对候选池按相似度排序,过滤低于 floor 的(提示非硬拒).
|
||||
|
||||
limit=5 上限:注册路径 registry 每条只存前 5 提示,防大池刷屏
|
||||
(codex review MEDIUM).
|
||||
"""
|
||||
out: list[dict] = []
|
||||
for name, expr in candidates.items():
|
||||
if not expression or not expr:
|
||||
continue
|
||||
try:
|
||||
r = similarity(expression, expr)
|
||||
except RecursionError:
|
||||
continue
|
||||
if r is not None and r >= floor:
|
||||
out.append({"name": name, "ratio": r})
|
||||
out.sort(key=lambda h: h["ratio"], reverse=True)
|
||||
return out[:limit]
|
||||
@@ -12,6 +12,7 @@ from __future__ import annotations
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import socket
|
||||
import sys
|
||||
from datetime import datetime
|
||||
@@ -68,6 +69,36 @@ def load_history_points(out_dir: str, exclude_file: str) -> dict[str, list[dict]
|
||||
return {f: [pts[m] for m in sorted(pts)] for f, pts in merged.items()}
|
||||
|
||||
|
||||
def factor_stats(points: list[dict]) -> dict:
|
||||
"""判定弹药聚合(2026-10-10 spec §4.2):全期合成 t/正 IC 占比/逐年.
|
||||
|
||||
口径:各月点 ic_mean(该批 12M 窗日均 IC 均值)的等权均值≈全史近似;
|
||||
合成 t=mean/(std(ddof=1)/√n),n<2 或 std=0 时如实 None(小样本判读弱
|
||||
的诚实注记在 spec,判定动作不因此自动化).
|
||||
"""
|
||||
ics = [p["ic_mean"] for p in points if p.get("ic_mean") is not None]
|
||||
if not ics:
|
||||
return {"icAll": None, "tAll": None, "positiveRatio": None, "byYear": {}}
|
||||
n = len(ics)
|
||||
mean_ic = sum(ics) / n
|
||||
t_all = None
|
||||
if n >= 2:
|
||||
var = sum((v - mean_ic) ** 2 for v in ics) / (n - 1)
|
||||
if var > 0:
|
||||
t_all = round(mean_ic / (var ** 0.5 / n ** 0.5), 4)
|
||||
by_year: dict[str, list[float]] = {}
|
||||
for p in points:
|
||||
m, v = str(p.get("month") or ""), p.get("ic_mean")
|
||||
# 月串须 YYYY-MM 才进逐年(codex review:非法串切前 4 字符会造出
|
||||
# "bad" 这类年键;ic 值聚合不含月串语义,照收不剔)
|
||||
if v is not None and re.fullmatch(r"\d{4}-\d{2}", m):
|
||||
by_year.setdefault(m[:4], []).append(v)
|
||||
return {"icAll": round(mean_ic, 6), "tAll": t_all,
|
||||
"positiveRatio": round(sum(1 for v in ics if v > 0) / n, 4),
|
||||
"byYear": {y: round(sum(v) / len(v), 6)
|
||||
for y, v in sorted(by_year.items())}}
|
||||
|
||||
|
||||
def build_report(registry: dict, current_points: dict[str, dict],
|
||||
history_points: dict[str, list[dict]], as_of: str,
|
||||
host: str) -> dict:
|
||||
@@ -80,8 +111,13 @@ def build_report(registry: dict, current_points: dict[str, dict],
|
||||
for name, entry in registry["factors"].items():
|
||||
cur = current_points.get(name) or {}
|
||||
hist = list(history_points.get(name, []))
|
||||
point = {"month": month_key, "t": cur.get("t")}
|
||||
if cur.get("t") is not None:
|
||||
# 当月点带 ic_mean(codex review:只带 t 会让 factor_stats 漏最新
|
||||
# 一期、首跑全空——t/ic_mean/count 同源于当月评估批)
|
||||
point = {"month": month_key, "t": cur.get("t"),
|
||||
"ic_mean": cur.get("ic_mean")}
|
||||
# 追加条件=t 或 ic_mean 任一非空(codex 复审 LOW:degenerate 月
|
||||
# sd==0 时 t=None 但 IC 有效,只看 t 会漏出 factor_stats 一期)
|
||||
if cur.get("t") is not None or cur.get("ic_mean") is not None:
|
||||
hist = [p for p in hist if p["month"] != month_key] + [point]
|
||||
monthly_points[name] = sorted(hist, key=lambda p: p["month"])
|
||||
if entry["status"] in ("promoted", "decaying"):
|
||||
@@ -112,6 +148,8 @@ def build_report(registry: dict, current_points: dict[str, dict],
|
||||
"observations": observations,
|
||||
"collective": collective_level(promoted_ts),
|
||||
"monthly_points": monthly_points,
|
||||
"factor_stats": {name: factor_stats(pts)
|
||||
for name, pts in monthly_points.items()},
|
||||
"unregistered": unregistered,
|
||||
"pre_gate": pre_gate_table(registry, monthly_points, current_points),
|
||||
"edge_cases": edge_cases,
|
||||
@@ -201,6 +239,22 @@ def render_markdown(report: dict) -> str:
|
||||
for k, v in report["observations"].items():
|
||||
rs = ";".join(v["reasons"]) if v["reasons"] else "新鲜且覆盖足"
|
||||
lines.append(f"- {v['status']} {k}: {rs}")
|
||||
stats = report.get("factor_stats") or {}
|
||||
if stats:
|
||||
lines.append("")
|
||||
lines.append("## 判定弹药(全期口径)")
|
||||
lines.append("")
|
||||
lines.append("| 因子 | 全期IC | 合成t | 正IC占比 | 逐年IC |")
|
||||
lines.append("|------|--------|-------|----------|--------|")
|
||||
for name in sorted(stats):
|
||||
s = stats[name]
|
||||
years = ", ".join(f"{y}:{v:.4f}" for y, v in s["byYear"].items()) or "—"
|
||||
|
||||
def _f(v, nd=4):
|
||||
return "—" if v is None else f"{v:.{nd}f}"
|
||||
|
||||
lines.append(f"| {name} | {_f(s['icAll'], 6)} | {_f(s['tAll'])} "
|
||||
f"| {_f(s['positiveRatio'])} | {years} |")
|
||||
c = report["collective"]
|
||||
lines.append(f"\n## 集体水位(决议 C:跌破=研判卡数据,权重冻结不动)\n"
|
||||
f"- 已晋级 n={c['n']},t 中位数={c['median']},地板={c['floor']},"
|
||||
|
||||
@@ -156,11 +156,30 @@ PY
|
||||
rc_replay=$?
|
||||
echo "=== $(date '+%F %T') stage5 verdict_replay exit=$rc_replay (派生件,不进 final) ==="
|
||||
|
||||
# stage6 challenger 三路对拍(影子实验,失败不阻月度链;产物落子目录
|
||||
# 判定层端点只读根层零影响)。宽表源=stage4 截面盘(12 源主 parquet,
|
||||
# daily_section 幂等 merge 累积)——月度批 values/ 只含注册表因子,
|
||||
# 不含 quant12 12 源(2026-10-10 现场核实 NAS values/monthly_2026-09)。
|
||||
"$DOCKER" run --rm --name sanguo-factor-chal --user "${UID_ADMIN}:${GID_ADMIN}" \
|
||||
--group-add "${GID_ADMINS}" --no-healthcheck --entrypoint python \
|
||||
-e HOME=/tmp -e MPLCONFIGDIR=/tmp/mpl \
|
||||
-v /volume1/stock:/volume1/stock \
|
||||
-v "$APP":/app:ro \
|
||||
-w /app \
|
||||
sanguo_vnpy_v2:lock-aligned \
|
||||
-m sanguo_factor.challenger_lgbm \
|
||||
--values-dir /volume1/stock/factor_cross_section \
|
||||
--as-of "$AS_OF" --out-dir "$OUT_DIR" \
|
||||
--vnpy-db /volume1/stock/sanguo_vnpy_v2/data_backup/quant_trading.db \
|
||||
--host nas
|
||||
rc_chal=$?
|
||||
echo "=== $(date '+%F %T') stage6 challenger exit=$rc_chal (影子,不进 final) ==="
|
||||
|
||||
# 真失败=②或③ rc=1(读件错);exit 2 是信号按 0 处理;④ 非零=真失败
|
||||
final=0
|
||||
[ "$rc_review" -eq 1 ] && final=1
|
||||
[ "$rc_gap" -eq 1 ] && final=1
|
||||
[ "$rc_attr" -ne 0 ] && final=1
|
||||
echo "=== $(date '+%F %T') factor-monthly done final_rc=$final (batch=$rc_batch review=$rc_review gap=$rc_gap attr=$rc_attr replay=$rc_replay) ==="
|
||||
echo "=== $(date '+%F %T') factor-monthly done final_rc=$final (batch=$rc_batch review=$rc_review gap=$rc_gap attr=$rc_attr replay=$rc_replay chal=$rc_chal) ==="
|
||||
exit "$final"
|
||||
} >> "$LOG" 2>&1
|
||||
|
||||
@@ -126,3 +126,63 @@ def test_graveyard_factor_list_endpoint(client, tmp_path):
|
||||
assert got["fa_old"]["cause"] == "IC 两窗反号(09-22 终判)"
|
||||
# 活跃因子不出现在墓园列表
|
||||
assert "fa_gross_margin" not in got
|
||||
|
||||
|
||||
def test_factor_detail_ic_stats(client, tmp_path, monkeypatch):
|
||||
"""detail 端点带出 factor_stats(全期弹药),无报告时如实 None."""
|
||||
import json as _json
|
||||
name, _ = _seed_factor_card(client, name="fa_x") # 注册活因子(assessable)
|
||||
rep_dir = tmp_path / "reports" / "factor_monthly"
|
||||
rep_dir.mkdir(parents=True)
|
||||
doc = {"as_of": "2026-09-30",
|
||||
"monthly_points": {"fa_x": [
|
||||
{"month": "2025-09", "ic_mean": 0.04, "t": 2.1, "count": 1},
|
||||
{"month": "2025-10", "ic_mean": 0.06, "t": 2.4, "count": 1},
|
||||
{"month": "2026-09", "ic_mean": 0.02, "t": 1.1, "count": 1}]},
|
||||
"factor_stats": {"fa_x": {"icAll": 0.04, "tAll": 1.7321,
|
||||
"positiveRatio": 1.0,
|
||||
"byYear": {"2025": 0.05, "2026": 0.02}}}}
|
||||
(rep_dir / "nas_2026-09-30.json").write_text(_json.dumps(doc), "utf-8")
|
||||
monkeypatch.setenv("SANGUO_FACTOR_MONTHLY_DIR", str(rep_dir))
|
||||
r = client.get(f"/api/v1/pipeline/factors/{name}/detail")
|
||||
assert r.status_code == 200
|
||||
got = r.json()["factor"]["icStats"]
|
||||
assert got["icAll"] == 0.04 and got["tAll"] == 1.7321
|
||||
assert got["byYear"]["2025"] == 0.05
|
||||
# 无报告时的反向用例:指到空目录,icStats 如实 None(不造数)
|
||||
empty = tmp_path / "reports_empty"
|
||||
empty.mkdir()
|
||||
monkeypatch.setenv("SANGUO_FACTOR_MONTHLY_DIR", str(empty))
|
||||
r2 = client.get(f"/api/v1/pipeline/factors/{name}/detail")
|
||||
assert r2.status_code == 200 and r2.json()["factor"]["icStats"] is None
|
||||
|
||||
|
||||
def test_factors_list_ic_full_t_from_stats(client, tmp_path, monkeypatch):
|
||||
"""codex review:列表端点 icFullT 从最新报告 factor_stats.tAll 接真
|
||||
(spec 承诺 factors 端点全史 t 升级);无报告仍如实 None."""
|
||||
import json as _json
|
||||
name, _ = _seed_factor_card(client, name="fa_x")
|
||||
rep_dir = tmp_path / "reports" / "factor_monthly"
|
||||
rep_dir.mkdir(parents=True)
|
||||
doc = {"as_of": "2026-09-30",
|
||||
"factor_stats": {"fa_x": {"icAll": 0.04, "tAll": 1.7321,
|
||||
"positiveRatio": 1.0,
|
||||
"byYear": {"2025": 0.05}}}}
|
||||
(rep_dir / "nas_2026-09-30.json").write_text(_json.dumps(doc), "utf-8")
|
||||
monkeypatch.setenv("SANGUO_FACTOR_MONTHLY_DIR", str(rep_dir))
|
||||
items = client.get("/api/v1/pipeline/factors").json()["items"]
|
||||
row = next(i for i in items if i["id"] == name)
|
||||
assert row["icFullT"] == 1.7321
|
||||
# tAll=None(样本不足)如实 null,不造数
|
||||
doc["factor_stats"]["fa_x"]["tAll"] = None
|
||||
(rep_dir / "nas_2026-09-30.json").write_text(_json.dumps(doc), "utf-8")
|
||||
items = client.get("/api/v1/pipeline/factors").json()["items"]
|
||||
row = next(i for i in items if i["id"] == name)
|
||||
assert row["icFullT"] is None
|
||||
# 无报告:零读不崩,如实 null
|
||||
empty = tmp_path / "reports_empty"
|
||||
empty.mkdir()
|
||||
monkeypatch.setenv("SANGUO_FACTOR_MONTHLY_DIR", str(empty))
|
||||
items = client.get("/api/v1/pipeline/factors").json()["items"]
|
||||
row = next(i for i in items if i["id"] == name)
|
||||
assert row["icFullT"] is None
|
||||
|
||||
@@ -295,6 +295,85 @@ class TestDecompose:
|
||||
from sanguo_api import routes_pipeline as rp
|
||||
assert isinstance(rp._DECOMPOSE_LOCK, type(threading.Lock()))
|
||||
assert isinstance(rp._GRADUATE_LOCK, type(threading.Lock()))
|
||||
def test_decompose_similarity_only_on_new_entry(self, client, fake_llm,
|
||||
tmp_path, monkeypatch):
|
||||
"""codex review 仅建条纪律:条目已存在时 upsert 不重写 similarity
|
||||
(强化形态——同名异假设 upsert 会拒,同名同假设幂等返回,此测后者)."""
|
||||
import yaml
|
||||
# 先建卡拿 hyp id,再预置 registry:fa_dup 绑同一假设(幂等重注册路径)
|
||||
hyp = _seed_card(client)
|
||||
reg_path = tmp_path / "r_sim2.yaml"
|
||||
monkeypatch.setenv("SANGUO_FACTOR_REGISTRY", str(reg_path))
|
||||
reg = {"factors": {
|
||||
"fa_old": {"name": "fa_old", "hypothesis": "H-1",
|
||||
"status": "assessable",
|
||||
"versions": [{"v": 1, "commit": "x",
|
||||
"params": {"expression":
|
||||
"ts_mean(close, 5) * turnover"},
|
||||
"effective_from": "2026-09-01"}]},
|
||||
"fa_dup": {"name": "fa_dup", "hypothesis": hyp,
|
||||
"status": "incubating",
|
||||
"versions": [{"v": 1, "commit": "x",
|
||||
"params": {"expression": "close + open"},
|
||||
"effective_from": "2026-09-01"}]}}}
|
||||
with open(reg_path, "w", encoding="utf-8") as f:
|
||||
yaml.safe_dump(reg, f, allow_unicode=True)
|
||||
# 绕开 LLM 与 name_taken 硬门,直接喂「已存在名」的 passed
|
||||
import sanguo_api.routes_pipeline as rp
|
||||
|
||||
async def fake_run(client_, card, **k):
|
||||
return {"passed": [{"name": "fa_dup",
|
||||
"expression": "ts_mean(close, 5) / volume",
|
||||
"description": "同窗均值除量",
|
||||
"justification": "j"}],
|
||||
"failed": [], "rounds": 1}
|
||||
|
||||
monkeypatch.setattr(rp, "_run_decompose", fake_run)
|
||||
assert _wait_job(client, hyp)["status"] == "completed"
|
||||
with open(reg_path, encoding="utf-8") as f:
|
||||
reg2 = yaml.safe_load(f)
|
||||
# 已存在条目:similarity 不写入(旧值 None 保持 None)
|
||||
assert "similarity" not in reg2["factors"]["fa_dup"]
|
||||
|
||||
def test_decompose_register_similarity_hint(self, client, fake_llm,
|
||||
tmp_path, monkeypatch):
|
||||
"""AST 防换皮提示(spec §4.2 2026-10-10):新因子注册时对既有池比对,
|
||||
疑似换皮(≥0.6)存 entry.similarity——硬门(dup_subtree)外的部分公共
|
||||
子树对仍可注册,提示非硬拒,判定人做."""
|
||||
import yaml
|
||||
reg_path = tmp_path / "r_sim.yaml"
|
||||
monkeypatch.setenv("SANGUO_FACTOR_REGISTRY", str(reg_path))
|
||||
# 既有池:fa_old 与候选共享 ts_mean(close,5) 子树(4 节点<硬门阈 8)
|
||||
reg = {"factors": {"fa_old": {
|
||||
"name": "fa_old", "hypothesis": "H-1", "status": "assessable",
|
||||
"versions": [{"v": 1, "commit": "x",
|
||||
"params": {"expression":
|
||||
"ts_mean(close, 5) * turnover"},
|
||||
"effective_from": "2026-09-01"}]}}}
|
||||
with open(reg_path, "w", encoding="utf-8") as f:
|
||||
yaml.safe_dump(reg, f, allow_unicode=True)
|
||||
_, queue = fake_llm
|
||||
queue.append({"factors": [
|
||||
{"name": "llm_twin", "expression": "ts_mean(close, 5) / volume",
|
||||
"description": "同窗均值除量", "justification": "j"}]})
|
||||
hyp = _seed_card(client)
|
||||
assert _wait_job(client, hyp)["status"] == "completed"
|
||||
with open(reg_path, encoding="utf-8") as f:
|
||||
reg2 = yaml.safe_load(f)
|
||||
entry = reg2["factors"]["llm_twin"]
|
||||
assert entry["similarity"] == [{"name": "fa_old", "ratio": 0.6667}]
|
||||
d = client.get("/api/v1/pipeline/factors/llm_twin/detail").json()
|
||||
assert d["factor"]["similarity"] == [{"name": "fa_old", "ratio": 0.6667}]
|
||||
# 列表端点同带出(工厂板换皮徽标数据源)
|
||||
items = client.get("/api/v1/pipeline/factors").json()["items"]
|
||||
row = next(i for i in items if i["id"] == "llm_twin")
|
||||
assert row["similarity"] == [{"name": "fa_old", "ratio": 0.6667}]
|
||||
|
||||
|
||||
|
||||
"""P3-7: save 端点 source 白名单+上限(此前零白名单零上限,任意串入库)."""
|
||||
|
||||
|
||||
|
||||
|
||||
class TestDecomposeJobs:
|
||||
@@ -517,10 +596,6 @@ class TestD7Visibility:
|
||||
assert client.get("/api/v1/pipeline/factors/nope/detail"
|
||||
).status_code == 404
|
||||
|
||||
|
||||
|
||||
"""P3-7: save 端点 source 白名单+上限(此前零白名单零上限,任意串入库)."""
|
||||
|
||||
class TestSaveSourceGuard:
|
||||
"""P3-7: save 端点 source 白名单+上限(此前零白名单零上限,任意串入库)."""
|
||||
|
||||
|
||||
@@ -0,0 +1,361 @@
|
||||
"""LGBM challenger 三路对拍(spec §4.4 2026-10-10):数据层."""
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
import pandas as pd
|
||||
import pytest
|
||||
|
||||
from sanguo_factor.challenger_lgbm import (
|
||||
build_label,
|
||||
cs_rank_norm,
|
||||
equal_weight_signal,
|
||||
icir_signal,
|
||||
load_pool,
|
||||
lgbm_walk_forward,
|
||||
run_challenge,
|
||||
score_signal,
|
||||
)
|
||||
|
||||
|
||||
def test_load_pool_twelve_sources():
|
||||
pool = load_pool()
|
||||
assert len(pool) == 12
|
||||
assert pool["vma_60"]["direction"] == "+"
|
||||
assert pool["vol_ma5"]["direction"] == "-"
|
||||
assert 0.0 < pool["alpha16"]["weight"] < 0.2
|
||||
|
||||
|
||||
def test_build_label_gap_proof():
|
||||
idx = pd.date_range("2025-01-01", periods=4, freq="D")
|
||||
close = pd.DataFrame({"a": [10.0, 11.0, 12.0, 13.0], "b": [20.0, 20.0, 21.0, 22.0]}, index=idx)
|
||||
lab = build_label(close)
|
||||
# label[t]=close[t+2]/close[t+1]-1:t0 行=12/11-1
|
||||
assert lab.iloc[0]["a"] == pytest.approx(12.0 / 11.0 - 1)
|
||||
# 末两行无可成交区间=NaN
|
||||
assert lab.iloc[-1].isna().all() and lab.iloc[-2].isna().all()
|
||||
|
||||
|
||||
def test_cs_rank_norm_nan_passthrough_and_centered():
|
||||
df = pd.DataFrame({"a": [1.0, 2.0, 3.0, np.nan],
|
||||
"b": [4.0, 3.0, 2.0, 1.0]})
|
||||
out = cs_rank_norm(df)
|
||||
# 行2截面 [a=3.0, b=2.0]: a 为最大, rank pct 1.0 - 0.5 = 0.5
|
||||
# (plan 原断言 0.75 与其自身注释"rank pct 1.0 - 0.5"矛盾, 按权威实现修正)
|
||||
assert out.loc[2, "a"] == pytest.approx(0.5)
|
||||
assert np.isnan(out.loc[3, "a"]) # NaN 透传不占截面
|
||||
row0 = out.loc[0, ["a", "b"]].tolist()
|
||||
assert row0 == [pytest.approx(0.0), pytest.approx(0.5)] # 截面最小/最大
|
||||
|
||||
|
||||
def _toy_values(idx, cols):
|
||||
rng = np.random.default_rng(7)
|
||||
return {n: pd.DataFrame(rng.normal(size=(len(idx), len(cols))), index=idx, columns=cols)
|
||||
for n in ("f_pos", "f_neg")}
|
||||
|
||||
|
||||
def test_equal_weight_direction_adjusted():
|
||||
idx = pd.date_range("2025-06-01", periods=3, freq="D")
|
||||
cols = ["a", "b"]
|
||||
values = {"f_pos": pd.DataFrame(1.0, index=idx, columns=cols),
|
||||
"f_neg": pd.DataFrame(1.0, index=idx, columns=cols)}
|
||||
pool = {"f_pos": {"direction": "+"}, "f_neg": {"direction": "-"}}
|
||||
sig = equal_weight_signal(values, pool)
|
||||
# +1 与 -1 等权平均=0
|
||||
assert (sig == 0.0).all().all()
|
||||
|
||||
|
||||
def test_icir_signal_uses_archive_weights():
|
||||
idx = pd.date_range("2025-06-01", periods=2, freq="D")
|
||||
cols = ["a"]
|
||||
values = {"f_pos": pd.DataFrame(2.0, index=idx, columns=cols),
|
||||
"f_neg": pd.DataFrame(2.0, index=idx, columns=cols)}
|
||||
pool = {"f_pos": {"direction": "+", "weight": 0.75},
|
||||
"f_neg": {"direction": "-", "weight": 0.25}}
|
||||
sig = icir_signal(values, pool)
|
||||
# 常值 2.0 单列截面的 CSRankNorm=0.5,期望=0.5*(+0.75-0.25)
|
||||
# (plan 原期望 2.0*(0.75-0.25) 未算入秩归一,按权威实现修正;
|
||||
# DataFrame 与 pytest.approx 标量不兼容,用 np.allclose)
|
||||
assert np.allclose(sig.values, 0.5 * (0.75 - 0.25))
|
||||
|
||||
|
||||
def test_lgbm_walk_forward_holdout_isolated():
|
||||
"""样外月绝不进训练(test 隔离铁律):给训练月与样外月截然不同的
|
||||
因子-收益关系,样外预测应反映训练期学到的关系而非记忆样外."""
|
||||
import pandas as pd
|
||||
from sanguo_factor.challenger_lgbm import build_label, lgbm_walk_forward
|
||||
idx = pd.date_range("2025-01-01", "2025-03-31", freq="B")
|
||||
cols = [f"s{i}" for i in range(5)]
|
||||
rng = np.random.default_rng(42)
|
||||
values = {"s0": pd.DataFrame(rng.normal(size=(len(idx), 5)), index=idx, columns=cols)}
|
||||
label = build_label(pd.DataFrame(100 + rng.normal(scale=0.5, size=(len(idx), 5)),
|
||||
index=idx, columns=cols))
|
||||
oos = idx[idx >= "2025-03-01"]
|
||||
# toy 数据(~215 样本)喂生产超参(lambda_l1=205+min_child_samples=100)会零分裂
|
||||
# →gain=0;走函数自带的 params 覆盖入口传轻量超参验证 split/gain 通路
|
||||
light_params = {"objective": "mse", "learning_rate": 0.1, "num_leaves": 4,
|
||||
"min_child_samples": 5, "lambda_l1": 0.0, "lambda_l2": 1.0,
|
||||
"seed": 42, "num_threads": 4, "verbose": -1}
|
||||
pred, imp = lgbm_walk_forward(values, label, last_month="2025-03", params=light_params)
|
||||
assert list(pred.index) == [d for d in oos if d in label.index and label.loc[d].notna().any()]
|
||||
assert set(imp.keys()) == {"s0"} and imp["s0"] > 0
|
||||
|
||||
|
||||
def test_score_signal_perfect_and_flat():
|
||||
# 5 列:实现有最小截面阈值(len(pair)<5 skip),plan 原稿 3 列进不了评分
|
||||
idx = pd.date_range("2025-03-03", periods=10, freq="B")
|
||||
cols = ["a", "b", "c", "d", "e"]
|
||||
sig = pd.DataFrame(np.linspace(-1, 1, 50).reshape(10, 5), index=idx, columns=cols)
|
||||
lab = sig * 1.0 # 完美信号
|
||||
s = score_signal(sig, lab)
|
||||
assert s["ic_mean"] == pytest.approx(1.0, abs=1e-6)
|
||||
assert s["q5q1"] > 0
|
||||
flat = pd.DataFrame(0.0, index=idx, columns=cols)
|
||||
s0 = score_signal(flat, lab)
|
||||
assert s0["days"] == 10 # 进评分的天数(IC=NaN 但天在)
|
||||
|
||||
|
||||
def test_run_challenge_end_to_end(tmp_path, monkeypatch):
|
||||
"""端到端:合成 2 因子×40 日数据落宽表 parquet+合成 close db 太重——
|
||||
本用例走 values_dir 真文件+vnpy_db 用 monkeypatch 替换拉取函数."""
|
||||
from sanguo_factor import challenger_lgbm as cl
|
||||
idx = pd.date_range("2025-01-01", "2025-03-31", freq="B")
|
||||
cols = ["a", "b"]
|
||||
rng = np.random.default_rng(3)
|
||||
values = {n: pd.DataFrame(rng.normal(size=(len(idx), 2)), index=idx, columns=cols)
|
||||
for n in ("s0", "s1")}
|
||||
vdir = tmp_path / "factor_values"
|
||||
vdir.mkdir()
|
||||
for n, v in values.items():
|
||||
v.to_parquet(vdir / f"{n}.parquet")
|
||||
close = pd.DataFrame(100 + np.cumsum(rng.normal(scale=0.4, size=(len(idx), 2)), axis=0),
|
||||
index=idx, columns=cols)
|
||||
# s0/s1 不在 quant12 档案内,plan 原稿缺此 patch 必 FileNotFoundError
|
||||
monkeypatch.setattr(cl, "load_pool", lambda: {
|
||||
"s0": {"direction": "+", "weight": 0.6},
|
||||
"s1": {"direction": "-", "weight": 0.4}})
|
||||
# 真实 loader(universe.load_universe_bars)必传 start/end,签名 4 参
|
||||
monkeypatch.setattr(cl, "_load_close_wide", lambda db, c, s, e: close)
|
||||
out = cl.run_challenge(str(vdir), "fake.db", "2025-03-31", str(tmp_path))
|
||||
doc = json.loads(Path(out).read_text())
|
||||
assert doc["as_of"] == "2025-03-31"
|
||||
assert set(doc["signals"]) == {"equal", "icir", "lgbm"}
|
||||
for k, v in doc["signals"].items():
|
||||
assert {"ic_mean", "icir", "q5q1", "days"} <= set(v)
|
||||
assert doc["pool"]["names"] == ["s0", "s1"]
|
||||
assert len(doc["feature_importance"]) == 2
|
||||
assert doc["split"]["oos_month"] == "2025-03"
|
||||
|
||||
|
||||
def test_cli_smoke(tmp_path, monkeypatch, capsys):
|
||||
from sanguo_factor import challenger_lgbm as cl
|
||||
idx = pd.date_range("2025-01-01", "2025-03-31", freq="B")
|
||||
close = pd.DataFrame(100 + np.cumsum(np.random.default_rng(1).normal(0.4, size=(len(idx), 2)), axis=0),
|
||||
index=idx, columns=["a", "b"])
|
||||
monkeypatch.setattr(cl, "_load_close_wide", lambda db, c, s, e: close)
|
||||
monkeypatch.setattr(cl, "load_pool", lambda: {
|
||||
"s0": {"direction": "+", "weight": 0.6},
|
||||
"s1": {"direction": "-", "weight": 0.4}})
|
||||
vdir = tmp_path / "vals"; vdir.mkdir()
|
||||
for n in ("s0", "s1"):
|
||||
pd.DataFrame(np.random.default_rng(2).normal(size=(len(idx), 2)),
|
||||
index=idx, columns=["a", "b"]).to_parquet(vdir / f"{n}.parquet")
|
||||
rc = cl.main(["--values-dir", str(vdir), "--as-of", "2025-03-31",
|
||||
"--out-dir", str(tmp_path), "--vnpy-db", "fake.db"])
|
||||
assert rc == 0
|
||||
assert (tmp_path / "challenger_lgbm" / "nas_2025-03-31.json").exists()
|
||||
|
||||
|
||||
# ———— codex 复审班次二修批(2026-10-10):H1/H7 embargo+训练窗裁拆 ————
|
||||
|
||||
def test_split_masks_embargo_and_floor():
|
||||
"""H-1/H→M-7:训练窗=[样外月前推 11 个自然月首,样外首日前 2 交易日).
|
||||
|
||||
embargo 2 交易的数学保证:训练样本 t 的 label 取 close[t+1],close[t+2]
|
||||
(build_label shift(-2)/(--1)),故 max(train_pos)+2 < first_oos_pos
|
||||
⇔ 样外月收盘价不参与任何训练样本的 label 计算.
|
||||
"""
|
||||
from sanguo_factor.challenger_lgbm import _split_masks
|
||||
idx = pd.date_range("2024-01-01", "2025-03-31", freq="B")
|
||||
train_mask, oos_mask = _split_masks(idx, "2025-03")
|
||||
oos_pos = np.flatnonzero(oos_mask)
|
||||
tr_pos = np.flatnonzero(train_mask)
|
||||
assert oos_pos[0] == int(np.argmax(oos_mask))
|
||||
# 样外月收盘绝不进训练 label:max 训练位 +2 严格早于样外首日
|
||||
assert tr_pos.max() + 2 < oos_pos[0]
|
||||
assert tr_pos.max() == oos_pos[0] - 3 # embargo=2 恰好剔够
|
||||
assert not train_mask[oos_pos].any() # 训练/样外不重叠
|
||||
# 训练下界=2025-03 前推 11 个自然月之首(2024-04-01)
|
||||
assert idx[tr_pos].min() >= pd.Timestamp("2024-04-01")
|
||||
# floor 之上只剔样外月+embargo 两天
|
||||
above = np.flatnonzero(np.asarray(idx >= pd.Timestamp("2024-04-01")))
|
||||
assert len(above) == len(tr_pos) + len(oos_pos) + 2
|
||||
# 无样外月:训练也空(防全量误训)
|
||||
t0, o0 = _split_masks(idx, "2026-01")
|
||||
assert not o0.any() and not t0.any()
|
||||
|
||||
|
||||
# ———— H-2:三路统一可评样本 mask ————
|
||||
|
||||
def test_unified_scoring_mask_same_sample_set(tmp_path, monkeypatch):
|
||||
"""某源缺 2 股:未统一时 equal/icir 按 6 股评分、lgbm 只剩 4 股(<5
|
||||
跳过)→days 分叉;统一 mask 后三路同一 date×stock 集合→days 一致."""
|
||||
from sanguo_factor import challenger_lgbm as cl
|
||||
idx = pd.date_range("2025-01-01", "2025-03-31", freq="B")
|
||||
cols = ["a", "b", "c", "d", "e", "f"]
|
||||
rng = np.random.default_rng(11)
|
||||
s0 = pd.DataFrame(rng.normal(size=(len(idx), 6)), index=idx, columns=cols)
|
||||
s0[["b", "c"]] = np.nan # 源 s0 缺 b/c 两股
|
||||
s1 = pd.DataFrame(rng.normal(size=(len(idx), 6)), index=idx, columns=cols)
|
||||
vdir = tmp_path / "vals"; vdir.mkdir()
|
||||
s0.to_parquet(vdir / "s0.parquet")
|
||||
s1.to_parquet(vdir / "s1.parquet")
|
||||
close = pd.DataFrame(100 + np.cumsum(rng.normal(scale=0.4, size=(len(idx), 6)), axis=0),
|
||||
index=idx, columns=cols)
|
||||
monkeypatch.setattr(cl, "load_pool", lambda: {
|
||||
"s0": {"direction": "+", "weight": 0.6},
|
||||
"s1": {"direction": "-", "weight": 0.4}})
|
||||
monkeypatch.setattr(cl, "_load_close_wide", lambda db, c, s, e: close)
|
||||
out = cl.run_challenge(str(vdir), "fake.db", "2025-03-31", str(tmp_path))
|
||||
doc = json.loads(Path(out).read_text())
|
||||
days = {k: v["days"] for k, v in doc["signals"].items()}
|
||||
assert len(set(days.values())) == 1 # 三路同一评分集
|
||||
assert set(days.values()) == {0} # 统一后每日常截面=4 股<5,全跳
|
||||
|
||||
|
||||
# ———— H-3/M-5:缺源与坏件走 skip 语义 ————
|
||||
|
||||
def test_missing_sources_skip_no_artifact(tmp_path, monkeypatch, capsys):
|
||||
from sanguo_factor import challenger_lgbm as cl
|
||||
idx = pd.date_range("2025-01-01", "2025-03-31", freq="B")
|
||||
vdir = tmp_path / "vals"; vdir.mkdir()
|
||||
pd.DataFrame(np.zeros((len(idx), 2)), index=idx,
|
||||
columns=["a", "b"]).to_parquet(vdir / "s0.parquet")
|
||||
monkeypatch.setattr(cl, "load_pool", lambda: {
|
||||
"s0": {"direction": "+", "weight": 0.5},
|
||||
"s1": {"direction": "+", "weight": 0.5}})
|
||||
rc = cl.main(["--values-dir", str(vdir), "--as-of", "2025-03-31",
|
||||
"--out-dir", str(tmp_path), "--vnpy-db", "fake.db"])
|
||||
assert rc == 0
|
||||
assert not (tmp_path / "challenger_lgbm").exists() # 不产半吊子件
|
||||
out = capsys.readouterr().out
|
||||
assert "missing sources" in out and "expected=2" in out and "loaded=1" in out
|
||||
assert "s1" in out
|
||||
|
||||
|
||||
def test_bad_parquet_structured_skip(tmp_path, monkeypatch, capsys):
|
||||
"""坏件(非 parquet 字节)/空件(0 行)结构化剔除不裸抛,汇入缺源 skip."""
|
||||
from sanguo_factor import challenger_lgbm as cl
|
||||
idx = pd.date_range("2025-01-01", "2025-03-31", freq="B")
|
||||
vdir = tmp_path / "vals"; vdir.mkdir()
|
||||
pd.DataFrame(np.zeros((len(idx), 2)), index=idx,
|
||||
columns=["a", "b"]).to_parquet(vdir / "s0.parquet")
|
||||
(vdir / "s1.parquet").write_bytes(b"definitely not a parquet")
|
||||
pd.DataFrame().to_parquet(vdir / "s2.parquet") # 0 行空件
|
||||
monkeypatch.setattr(cl, "load_pool", lambda: {
|
||||
n: {"direction": "+", "weight": 1.0 / 3} for n in ("s0", "s1", "s2")})
|
||||
rc = cl.main(["--values-dir", str(vdir), "--as-of", "2025-03-31",
|
||||
"--out-dir", str(tmp_path), "--vnpy-db", "fake.db"])
|
||||
assert rc == 0
|
||||
assert not (tmp_path / "challenger_lgbm").exists()
|
||||
out = capsys.readouterr().out
|
||||
assert "expected=3" in out and "loaded=1" in out
|
||||
assert "s1" in out and "s2" in out
|
||||
|
||||
|
||||
# ———— H-4:_load_close_wide 真实路径(monkeypatch universe.load_universe_bars) ————
|
||||
|
||||
def test_load_close_wide_real_path(monkeypatch):
|
||||
import polars as pl
|
||||
import sanguo_factor.universe as uni
|
||||
from sanguo_factor import challenger_lgbm as cl
|
||||
calls = {}
|
||||
|
||||
def fake_loader(db, start, end, symbols=None, limit=None):
|
||||
calls.update(db=db, start=start, end=end, symbols=list(symbols or []))
|
||||
return pl.DataFrame({
|
||||
"vt_symbol": ["a.SSE", "a.SSE", "b.SSE", "b.SSE"],
|
||||
"datetime": [pd.Timestamp("2025-01-02"), pd.Timestamp("2025-01-03")] * 2,
|
||||
"close": [10.0, 11.0, 20.0, 21.0],
|
||||
})
|
||||
|
||||
monkeypatch.setattr(uni, "load_universe_bars", fake_loader)
|
||||
wide = cl._load_close_wide("fake.db", pd.Index(["a.SSE", "b.SSE"]),
|
||||
"2025-01-01", "2025-01-31")
|
||||
assert calls["db"] == "fake.db"
|
||||
assert calls["start"] == "2025-01-01" and calls["end"] == "2025-01-31"
|
||||
assert calls["symbols"] == ["a.SSE", "b.SSE"]
|
||||
assert isinstance(wide.index, pd.DatetimeIndex)
|
||||
assert list(wide.columns) == ["a.SSE", "b.SSE"]
|
||||
assert wide.loc[pd.Timestamp("2025-01-03"), "a.SSE"] == pytest.approx(11.0)
|
||||
# 反向:loader 空帧 → 空宽表不炸
|
||||
monkeypatch.setattr(uni, "load_universe_bars", lambda *a, **k: pl.DataFrame(
|
||||
schema={"vt_symbol": pl.Utf8, "datetime": pl.Datetime, "close": pl.Float64}))
|
||||
wide0 = cl._load_close_wide("fake.db", pd.Index([]), "2025-01-01", "2025-01-31")
|
||||
assert wide0.empty
|
||||
|
||||
|
||||
# ———— M-6:cols_ref=池内 union ————
|
||||
|
||||
def test_cols_ref_is_union_of_pool_columns(tmp_path, monkeypatch):
|
||||
from sanguo_factor import challenger_lgbm as cl
|
||||
idx = pd.date_range("2025-01-01", "2025-03-31", freq="B")
|
||||
rng = np.random.default_rng(5)
|
||||
s0 = pd.DataFrame(rng.normal(size=(len(idx), 2)), index=idx, columns=["a", "b"])
|
||||
s1 = pd.DataFrame(rng.normal(size=(len(idx), 2)), index=idx, columns=["b", "c"])
|
||||
vdir = tmp_path / "vals"; vdir.mkdir()
|
||||
s0.to_parquet(vdir / "s0.parquet")
|
||||
s1.to_parquet(vdir / "s1.parquet")
|
||||
captured = {}
|
||||
|
||||
def fake_close(db, cols_ref, start, end):
|
||||
captured["cols"] = list(cols_ref)
|
||||
return pd.DataFrame(100 + rng.normal(scale=0.4, size=(len(idx), len(cols_ref))),
|
||||
index=idx, columns=list(cols_ref))
|
||||
|
||||
monkeypatch.setattr(cl, "load_pool", lambda: {
|
||||
"s0": {"direction": "+", "weight": 0.5},
|
||||
"s1": {"direction": "-", "weight": 0.5}})
|
||||
monkeypatch.setattr(cl, "_load_close_wide", fake_close)
|
||||
cl.run_challenge(str(vdir), "fake.db", "2025-03-31", str(tmp_path))
|
||||
assert captured["cols"] == ["a", "b", "c"]
|
||||
|
||||
|
||||
# ———— M-8:池冗余按日循环 ————
|
||||
|
||||
def test_redundancy_flags_per_day_unit():
|
||||
from sanguo_factor.challenger_lgbm import _redundancy_flags
|
||||
idx = pd.date_range("2025-01-01", periods=60, freq="B")
|
||||
cols = [f"s{i}" for i in range(20)]
|
||||
rng = np.random.default_rng(9)
|
||||
x = pd.DataFrame(rng.normal(size=(60, 20)), index=idx, columns=cols)
|
||||
z = pd.DataFrame(rng.normal(size=(60, 20)), index=idx, columns=cols)
|
||||
flagged = _redundancy_flags({"x": x, "y": x.copy(), "z": z})
|
||||
pairs = {(f["a"], f["b"]) for f in flagged}
|
||||
assert ("x", "y") in pairs # 完全同源必标
|
||||
assert ("x", "z") not in pairs and ("y", "z") not in pairs # 独立源不标
|
||||
xy = next(f for f in flagged if (f["a"], f["b"]) == ("x", "y"))
|
||||
assert xy["corr"] == pytest.approx(1.0, abs=1e-6)
|
||||
|
||||
|
||||
def test_unsorted_or_dupe_index_parquet_rejected(tmp_path, monkeypatch, capsys):
|
||||
"""codex 复审尾批:乱序/重复日期件 structured reject 并入缺源 skip.
|
||||
|
||||
embargo 的 shift(-2) 数学以位置序为地基,乱序/重复日期件必须拒."""
|
||||
from sanguo_factor import challenger_lgbm as cl
|
||||
idx = pd.date_range("2025-01-01", "2025-03-31", freq="B")
|
||||
vdir = tmp_path / "vals"; vdir.mkdir()
|
||||
good = pd.DataFrame(np.zeros((len(idx), 2)), index=idx, columns=["a", "b"])
|
||||
good.to_parquet(vdir / "s0.parquet")
|
||||
good.iloc[::-1].to_parquet(vdir / "s1.parquet") # 乱序:日期倒排
|
||||
pd.concat([good.iloc[:1], good]).to_parquet(vdir / "s2.parquet") # 重复首日
|
||||
monkeypatch.setattr(cl, "load_pool", lambda: {
|
||||
n: {"direction": "+", "weight": 1.0 / 3} for n in ("s0", "s1", "s2")})
|
||||
rc = cl.main(["--values-dir", str(vdir), "--as-of", "2025-03-31",
|
||||
"--out-dir", str(tmp_path), "--vnpy-db", "fake.db"])
|
||||
assert rc == 0
|
||||
assert not (tmp_path / "challenger_lgbm").exists()
|
||||
out = capsys.readouterr().out
|
||||
assert "expected=3" in out and "loaded=1" in out
|
||||
assert "s1" in out and "s2" in out
|
||||
assert "乱序" in out
|
||||
@@ -0,0 +1,89 @@
|
||||
# tests/factor/test_expression_match.py
|
||||
"""AST 防换皮(spec §4.2 2026-10-10):交换律归一化+最大公共子树占比."""
|
||||
from sanguo_factor.expression_match import similarity, top_similar
|
||||
|
||||
|
||||
def test_identical_and_commutative():
|
||||
assert similarity("close/open", "close/open") == 1.0
|
||||
# 交换律:a+b ≡ b+a(归一化后同构)
|
||||
assert similarity("close + open", "open + close") == 1.0
|
||||
assert similarity("ts_mean(close, 5) * volume",
|
||||
"volume * ts_mean(close, 5)") == 1.0
|
||||
|
||||
|
||||
def test_partial_common_subtree():
|
||||
# 节点口径(Name/Constant/Call 各计 1,算子与 ctx 为元数据不计):
|
||||
# 公共子树 ts_mean(close,5)=4 节点(Call+func+close+5);
|
||||
# 整树=BinOp+Call+volume(turnover)=6 节点;ratio=4/6≈0.6667
|
||||
r = similarity("ts_mean(close, 5) / volume",
|
||||
"ts_mean(close, 5) * turnover")
|
||||
assert r is not None and 0.6 < r < 1.0
|
||||
|
||||
|
||||
def test_no_common_and_invalid():
|
||||
assert similarity("close / open", "volume * turnover") == 0.0
|
||||
assert similarity("close +++", "close/open") is None # 解析失败如实 None
|
||||
assert similarity("", "close") is None
|
||||
|
||||
|
||||
def test_top_similar_floor():
|
||||
cands = {"fa_a": "close + open", "fa_b": "open + close",
|
||||
"fa_c": "volume * turnover"}
|
||||
hits = top_similar("close + open", cands, floor=0.6)
|
||||
assert [h["name"] for h in hits] == ["fa_a", "fa_b"]
|
||||
assert all(h["ratio"] >= 0.6 for h in hits)
|
||||
# 空表达式/坏 candidates 不炸
|
||||
assert top_similar("close", {"fa_x": ""}) == []
|
||||
|
||||
|
||||
# —— codex review 修批(2026-10-10):UnaryOp/keywords/limit/深树防护 ——
|
||||
def test_unaryop_distinct_operands_not_similar():
|
||||
"""-close vs -open:UnaryOp 显式编码,不再落 ? 兜底桶互相撞签."""
|
||||
assert similarity("-close", "-open") == 0.0
|
||||
assert similarity("-ts_mean(close,5)", "-ts_mean(open,99)") == 0.0
|
||||
# 同构 UnaryOp 仍 1.0
|
||||
assert similarity("-close", "-close") == 1.0
|
||||
# containment 语义:close 整个是 -close 的子树,较小表达式全包含→1.0
|
||||
# (口径=公共子树/较小表达式节点数,非对称同构判定)
|
||||
assert similarity("-close", "close") == 1.0
|
||||
|
||||
|
||||
def test_call_keywords_encoded():
|
||||
"""Call 关键字参数须进签名:kw 值不同→不再 1.0(共享 close 操作数
|
||||
仅余 1/5 地板噪声,低于提示阈 0.6)."""
|
||||
r = similarity("ts_mean(close, window=5)", "ts_mean(close, window=99)")
|
||||
assert r is not None and r < 0.6
|
||||
# 同 kwargs 同构
|
||||
assert similarity("f(close, w=5)", "f(close, w=5)") == 1.0
|
||||
# kwargs 顺序无关(按名排序后拼接)
|
||||
assert similarity("f(a=1, b=2)", "f(b=2, a=1)") == 1.0
|
||||
# kwargs 有无也是结构差异:无 kw vs 有 kw 不撞签(仅共享 close,1/3)
|
||||
r2 = similarity("f(close)", "f(close, w=5)")
|
||||
assert r2 is not None and r2 < 0.6
|
||||
|
||||
|
||||
def test_nested_call_common_subtree():
|
||||
"""嵌套 Call:全同→1;算子名也不同→0;仅共享外层算子名→占比 1/6 低于提示阈."""
|
||||
assert similarity("cs_rank(ts_mean(close, 5))",
|
||||
"cs_rank(ts_mean(close, 5))") == 1.0
|
||||
assert similarity("cs_rank(ts_mean(close, 5))",
|
||||
"zscore(ts_mean(open, 9))") == 0.0
|
||||
r = similarity("cs_rank(ts_mean(close, 5))", "cs_rank(ts_mean(open, 9))")
|
||||
assert r is not None and r < 0.6
|
||||
|
||||
|
||||
def test_top_similar_limit_and_empty():
|
||||
"""limit=5 上限(注册路径 registry 只存前 5)+空候选不炸."""
|
||||
cands = {f"fa_{i}": "close + open" for i in range(7)}
|
||||
hits = top_similar("close + open", cands)
|
||||
assert len(hits) == 5
|
||||
assert top_similar("close", {}) == []
|
||||
assert top_similar("", {"fa_a": "close"}) == []
|
||||
|
||||
|
||||
def test_deep_expression_no_crash():
|
||||
"""深表达式递归爆栈不 raise,如实 None(factor_guard _DEPTH_MAX=12 挡真实
|
||||
注册,此处为纯防护——RecursionError 捕获返回 None)."""
|
||||
deep = "close" + " + close" * 3000
|
||||
assert similarity(deep, deep) is None
|
||||
assert top_similar(deep, {"fa_a": deep}) == []
|
||||
@@ -1,6 +1,7 @@
|
||||
# tests/factor/test_monthly_review.py
|
||||
"""月度批评 TDD——spec §4.2 决议 H: 双机各跑,数据截止对齐自然月末."""
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
|
||||
import pytest
|
||||
@@ -9,6 +10,7 @@ from sanguo_factor import eval_store
|
||||
from sanguo_factor.monthly_review import (
|
||||
build_report,
|
||||
extract_points,
|
||||
factor_stats,
|
||||
load_history_points,
|
||||
main,
|
||||
pick_eval_run,
|
||||
@@ -103,8 +105,9 @@ def test_build_report_verdict_and_collective(eval_db, registry_yaml):
|
||||
assert rep["collective"]["is_collective_decay"] is False
|
||||
# 观察面: s01_buzz 无批记录 → red
|
||||
assert rep["observations"]["s01_buzz"]["status"] == "red"
|
||||
# 当月点已存入 monthly_points(下月续算原料)
|
||||
assert rep["monthly_points"]["pledge_net_chg"][-1] == {"month": "2026-09", "t": 1.2}
|
||||
# 当月点已存入 monthly_points(下月续算原料;codex review 后带 ic_mean)
|
||||
assert rep["monthly_points"]["pledge_net_chg"][-1] == {
|
||||
"month": "2026-09", "t": 1.2, "ic_mean": 0.028}
|
||||
|
||||
|
||||
def test_history_points_extend_streak(eval_db, registry_yaml, tmp_path):
|
||||
@@ -253,3 +256,81 @@ def test_pre_gate_threshold_follows_promotion_bar_constant(monkeypatch):
|
||||
pg = mr.pre_gate_table(reg, {"fa_a": [{"month": "2026-09", "t": 2.2}]},
|
||||
{"fa_a": {"t": 2.2, "ic_mean": 0.05, "count": 200}})
|
||||
assert pg["fa_a"]["flag"] == "REVIEW" # 2.2 < 2.5(常量)→REVIEW;硬编码 2.0 会 PASS
|
||||
|
||||
|
||||
# —— 判定弹药三件(2026-10-10 spec §4.2): 全期合成t/正IC占比/逐年IC ——
|
||||
def _ammo_pts(*pairs):
|
||||
return [{"month": m, "ic_mean": v, "t": None, "count": 1} for m, v in pairs]
|
||||
|
||||
|
||||
def test_factor_stats_basic():
|
||||
s = factor_stats(_ammo_pts(("2025-09", 0.04), ("2025-10", 0.06), ("2026-09", 0.02)))
|
||||
assert s["icAll"] == round((0.04 + 0.06 + 0.02) / 3, 6)
|
||||
assert s["positiveRatio"] == 1.0
|
||||
assert s["byYear"] == {"2025": round((0.04 + 0.06) / 2, 6), "2026": 0.02}
|
||||
mean = 0.04
|
||||
var = ((0.0) + (0.02 ** 2) + ((-0.02) ** 2)) / 2 # ddof=1(离差平方)
|
||||
t_all = mean / (math.sqrt(var) / math.sqrt(3))
|
||||
assert s["tAll"] == round(t_all, 4)
|
||||
|
||||
|
||||
def test_factor_stats_negative_and_short():
|
||||
s = factor_stats(_ammo_pts(("2025-09", 0.05), ("2025-10", -0.01)))
|
||||
assert s["positiveRatio"] == 0.5
|
||||
assert s["tAll"] is not None
|
||||
single = factor_stats(_ammo_pts(("2025-09", 0.05)))
|
||||
assert single["tAll"] is None # n<2 无合成 t
|
||||
assert single["positiveRatio"] == 1.0
|
||||
|
||||
|
||||
def test_factor_stats_empty_and_dirty():
|
||||
assert factor_stats([]) == {"icAll": None, "tAll": None,
|
||||
"positiveRatio": None, "byYear": {}}
|
||||
dirty = [{"month": "2025-09", "ic_mean": None, "t": None},
|
||||
{"month": "", "ic_mean": 0.1, "t": None}]
|
||||
s = factor_stats(dirty)
|
||||
assert s["icAll"] == 0.1
|
||||
assert s["byYear"] == {} # 空 month 不进逐年
|
||||
|
||||
|
||||
# —— codex review 修批(2026-10-10):point 带 ic_mean/非法 month 校验/退化样本 ——
|
||||
def test_factor_stats_degenerate_samples():
|
||||
"""n=2 同值→方差 0→tAll 如实 None(不除零不造 0)."""
|
||||
s = factor_stats(_ammo_pts(("2025-09", 0.05), ("2025-10", 0.05)))
|
||||
assert s["icAll"] == 0.05 and s["tAll"] is None
|
||||
assert s["positiveRatio"] == 1.0
|
||||
|
||||
|
||||
def test_factor_stats_invalid_month_not_in_by_year():
|
||||
"""非法 month(bad-2025)不切前 4 字符进 byYear;ic 值聚合不受月串影响."""
|
||||
pts = [{"month": "bad-2025", "ic_mean": 0.1, "t": None},
|
||||
{"month": "2025-09", "ic_mean": 0.04, "t": None}]
|
||||
s = factor_stats(pts)
|
||||
assert s["byYear"] == {"2025": 0.04} # bad-2025 整条剔出逐年
|
||||
assert s["icAll"] == round((0.1 + 0.04) / 2, 6) # 值聚合照收
|
||||
|
||||
|
||||
def test_build_report_current_point_carries_ic_mean():
|
||||
"""当月 point 须带 ic_mean——否则 factor_stats 漏最新一期、首跑全空."""
|
||||
reg = {"factors": {"fa_x": {"name": "fa_x", "hypothesis": "h",
|
||||
"status": "assessable"}}}
|
||||
rep = build_report(reg, {"fa_x": {"t": 2.0, "ic_mean": 0.05, "count": 100}},
|
||||
history_points={}, as_of="2026-09-30", host="t")
|
||||
assert rep["monthly_points"]["fa_x"][-1]["ic_mean"] == 0.05
|
||||
s = rep["factor_stats"]["fa_x"]
|
||||
assert s["icAll"] == 0.05 # 首跑单点即有弹药
|
||||
assert s["positiveRatio"] == 1.0
|
||||
assert s["tAll"] is None # n=1 无合成 t,如实
|
||||
|
||||
|
||||
# —— codex 复审 LOW:degenerate 当月(t=None 但 IC 有效)不漏出 factor_stats ——
|
||||
def test_build_report_degenerate_current_point_kept():
|
||||
"""当月 sd==0 时 t=None 但 ic_mean 有效→当月点须进 monthly_points
|
||||
与 factor_stats(追加条件=t 或 ic_mean 任一非空)."""
|
||||
reg = {"factors": {"fa_x": {"name": "fa_x", "hypothesis": "h",
|
||||
"status": "assessable"}}}
|
||||
rep = build_report(reg, {"fa_x": {"t": None, "ic_mean": 0.03, "count": 100}},
|
||||
history_points={}, as_of="2026-09-30", host="t")
|
||||
assert rep["monthly_points"]["fa_x"] == [
|
||||
{"month": "2026-09", "t": None, "ic_mean": 0.03}]
|
||||
assert rep["factor_stats"]["fa_x"]["icAll"] == 0.03
|
||||
|
||||
Reference in New Issue
Block a user