# 中证1000/2000 历史成份股补全 实施计划 > **For agentic workers:** REQUIRED SUB-SKILL: superpowers:executing-plans。Steps 用 checkbox `- [ ]` 跟踪。 **Goal:** 把中证1000(000852)/中证2000(932000)从"纯当前快照"补成"治幸存者偏差的全集"(含被踢出的股票),接入 `constituent_unified`,并部署定期更新 schtask。 **Architecture:** csindex 官方公告 JSON 接口(`queryAnnouncementByVo` + `queryAnnouncementById`)抓调整公告 → 解析附件 PDF/xlsx 的调入/调出名单 → 聚合成"曾经入选集"(全集型,非时点型)→ 入 `constituent_unified`,`in_current`=当前快照、`was_removed`=曾经入选−当前。 **Tech Stack:** Python3 + pandas + openpyxl + pdfplumber + sqlite3 + PowerShell schtask ## 诊断(已实证,2026-07-23) 现状 `constituent_unified`(VPS quant_trading.db): - `000852`: total=1000, in_current=1000, **was_removed=0**(纯快照,未治偏差) - `932000`: total=2000, in_current=2000, **was_removed=0**(纯快照) - 对比 `000300`: total=940, in_current=300, was_removed=640(已治偏差) **三处断点:** 1. **932000 launch xlsx 解析 bug**:`parse_csindex_announce.py:529` 取 `row[0]`(=指数代码 932000),应为 `row[3]`(证券代码)。→ 产出 distinct=1(2000 行全是 932000)。xlsx 实证 6 列:`指数代码/指数简称/指数英文简称/证券代码/证券中文简称/证券英文名称`。 2. **000852 公告覆盖不全**:`filter_csi1000_notices`(162 行)用 `theme='指数调样'+title 含'中证1000'` 过滤,只拿 28 份(2018-07 起)。调查实证:列表 API payload 加 `indexCode:'000852'` 能拿 **96 条**(45 调样),可回溯到 **2014 发布期**(早期 HTML 表格,2018+ PDF/xlsx)。 3. **migrate 没接 announce_union**:`migrate_constituent.py:118-131` 只读 `_snapshot.parquet`,没读 `_announce_union.parquet`。→ 1220 个治偏差集白产了。路径也对不上(parse 在 Mac 产 announce_union,migrate 读 VPS HIST,没同步)。 ## 关键简化 `constituent_unified` 是**全集型**(300/500/50 = baostock 988 时点聚合成 in_current/was_removed),**不是时点型**。所以: - **不需要**反向回溯引擎(生效日边界、逐时点 asof join) - 只要"曾经入选集"= 所有公告 add 记录 ∪ initial ∪ current 的 distinct code - `in_current` = akshare 当前快照(权威),`was_removed` = 曾经入选 − 当前 调查 agent 提的"生效日≠公告日"等坑是**时点型**需求才需要,本计划(全集型)不涉及。 ## Global Constraints(spec 铁律) - baostock 单进程单登录不并发(本计划不碰 baostock,无冲突) - 直连不走代理:`unset http_proxy https_proxy all_proxy`(脚本已内置) - 单线程限速:csindex 接口 sleep 1.0~1.5s - staging→验证→合并,不直接写主库(migrate 走 staging→merge 两步,已幂等) - provider 读 VPS 本地,不调 online(本计划是采集层,可调 csindex) - commit message 无 Co-Authored-By --- ### Task 1(#30):修 parse_csindex_announce.py 两处 **Files:** - Modify: `scripts/data_platform/parse_csindex_announce.py:526-538`(932000 launch xlsx 列索引) - Modify: `scripts/data_platform/parse_csindex_announce.py:120-182`(000852 列表搜索用 indexCode) **改动 1a — 932000 launch xlsx 列索引(:526-538):** 现:`code = _norm_code(row[0])`, `name = str(row[1])`。改为按 header 定位列(稳健),或直接 `code=row[3]`, `name=row[4]`。推荐 header 定位: ```python header = rows[0] # 找"证券代码"和"证券中文简称"列(中英文混合 header) code_idx = next((i for i,h in enumerate(header) if h and "证券代码" in str(h)), 3) name_idx = next((i for i,h in enumerate(header) if h and "证券中文简称" in str(h)), 4) for row in rows[1:]: code = _norm_code(row[code_idx] if len(row)>code_idx else None) name = str(row[name_idx]).strip() if len(row)>name_idx and row[name_idx] else "" ``` **改动 1b — 000852 列表搜索用 indexCode(:120-182):** 现 `fetch_all_notices` 拉全量再 `filter_csi1000_notices` title 过滤。改为:对 000852 用 `indexCode` payload 直接搜: ```python payload = {"lang":"cn","classlist":[],"indexlist":[], "indexCode":"000852", # ← 新增,直接按指数搜 "page":{"desc":"","key":"","page":page,"rows":100}, "related_topics":[],"typelist":[]} ``` 保留旧 filter 作兜底(标题含中证1000+调整)。合并 indexCode 命中 ∪ 已知 REGULAR/TEMP_IDS 去重。932000 走全局 `related_topics:["index_rebalance"]` + PDF grep "中证2000" section(parse_pdf_adjustments 已支持 target_section)。 **验证探针:** ```bash python3 scripts/data_platform/parse_csindex_announce.py --only 1000 # 期望:filtered CSI 1000 公告 ≥ 40 条(原 28),date 范围早于 2018-07 python3 scripts/data_platform/parse_csindex_announce.py --only 2000 # 期望:932000_announce_union.parquet distinct codes ≈ 2000(原 bug=1) ``` - [ ] Step 1: 改 932000 launch xlsx 列索引(header 定位) - [ ] Step 2: 改 000852 列表搜索(indexCode payload + 932000 related_topics) - [ ] Step 3: Mac 重跑 `--only 1000` + `--only 2000`,验证探针 - [ ] Step 4: commit --- ### Task 2(#31):改 migrate_constituent.py 接 announce_union 聚合全集 **Files:** - Modify: `scripts/data_platform/migrate_constituent.py:118-131`(加读 announce_union) - Test: `tests/portfolio/test_migrate_announce_union.py`(新建,TDD) **聚合逻辑(全集型):** ```python # 读 000852_announce_union.parquet + 932000_announce_union.parquet # announce_union schema: updateDate/index_code/code/code_name/adjust_type(add|remove|current|initial|current)/notice_id/source # 全集聚合: for idx in ['000852','932000']: ann = read(f"{idx}_announce_union.parquet") snap = read(f"{idx}_snapshot.parquet") # akshare 当前快照,权威 in_current current_codes = set(snap['code']) # 当前在册 ever_codes = set(ann['code']) | current_codes # 曾经入选(所有 add/initial + current) # 产出:ever_codes 每只一行 # in_current = code in current_codes # was_removed = code not in current_codes(曾入选已踢) # source = 'csindex_announce' ``` schema 对齐:`index_code/code/code_name/source/in_current/was_removed`。`code_name` 取 announce_union 或 snapshot 的(优先 snapshot 当前名)。 **合并进 staging:** 现有 `all_df = pd.concat([pool, df_deep, df_snap])`(:134)→ 把 000852/932000 的 announce_union 全集**替换** df_snap 里的 000852/932000 快照行(快照并入 announce 全集的 in_current),其他指数不动。 **TDD 测试(tests/portfolio/test_migrate_announce_union.py):** - test announce_union 聚合:given announce(add A,B + remove C) + snapshot(current A,B,D),assert ever={A,B,C,D}, in_current={A,B,D}, was_removed={C} - test 000852 distinct > 1000(治偏差证据) - test 932000 distinct ≈ 2000(launch 修复) - test 幂等(跑两次结果一致) - [ ] Step 1: 写聚合测试(RED) - [ ] Step 2: 改 migrate 加 announce_union 聚合(GREEN) - [ ] Step 3: 测试通过 - [ ] Step 4: commit --- ### Task 3(#32):重跑→同步VPS→migrate→merge→验证 **Files:** 无新文件(运行现有 pipeline) - [ ] Step 1: Mac 重跑 parse_csindex_announce.py --only both → 新 announce_union - [ ] Step 2: scp 000852_announce_union.parquet + 932000_announce_union.parquet 到 VPS `C:\sanguo_vnpy_v2\data\index_const_hist\` - [ ] Step 3: rsync 改后的 migrate_constituent.py 到 VPS - [ ] Step 4: VPS 跑 migrate_constituent.py(SANGUO_DB 指向 quant_trading.db)→ merge_constituent.py - [ ] Step 5: 验证(见下) **验证标准(VPS 查 constituent_unified):** ```sql SELECT index_code, COUNT(*), SUM(in_current), SUM(was_removed) FROM constituent_unified WHERE index_code IN ('000852','932000') GROUP BY index_code; ``` - 000852: total > 1000(曾经入选 ~1200+), in_current=1000, **was_removed > 0**(治偏差) - 932000: total ≈ 2000+, in_current=当前快照数, was_removed ≥ 0(launch ∪ current,中间调整无记录则 was_removed=0 可接受) - 抽样:挑一只 known 被踢股(如 announce_union 里 remove 类型)→ constituent_unified 该 code was_removed=1 - 回归:300/500/50/深证 行数不变(没误伤) --- ### Task 4(#33):定期 schtask 方案+部署 **Files:** - Create: `scripts/data_platform/csindex_constituent_wrapper.ps1` - Create: `scripts/data_platform/register_csindex_schtasks.ps1` **schtask 设计:** - 名:`sanguo-csindex-constituent` - 频率:**每月 16 号 + 6月/12月定调后额外**(中证1000 定期调整 6月/12月,临时调整不定期 → 月度抓足够,缓存增量) - 时间:**20:30**(避开 baostock 18:05/xt 18:40/akshare 19:00-19:50 窗口) - 流程:parse_csindex_announce.py --refresh-list(抓新公告)→ 同步 announce_union 已在本机 → migrate → merge - 幂等:migrate/merge 已 DROP+CREATE 可重跑;parse 有 notice cache 增量 **wrapper ps1(仿 bs_eod_wrapper.ps1 风格):** unset proxy → Set-Location → timestamped log → python parse + migrate + merge → exit code - [ ] Step 1: 写 wrapper ps1 + register ps1 - [ ] Step 2: VPS 部署 + schtasks /create /ru SYSTEM /rl HIGHEST - [ ] Step 3: 手动触发一次验证(schtasks /run) - [ ] Step 4: commit + 同步安装目录 --- ### Task 5(#34):更新 memory **Files:** - Update: memory `data-fusion-design-finalized.md`(推翻 000852/932000 "永久 gap") - Update: memory `static_data_gaps_design.md`(中证1000/2000 gap 关闭) - Update: `MEMORY.md` 索引 **记:** csindex 公告 JSON 接口路推翻"永久 gap";000852 全集入库(曾经入选 1200+);932000 launch xlsx 列 bug 修复;全集型简化洞察(不需回溯引擎);定期 schtask;调查 agent 实证的 96 公告/45 调样/回溯到 2014。 - [ ] Step 1: 更新 3 个 memory 文件 - [ ] Step 2: MEMORY.md 索引行 --- ## Self-Review - spec 覆盖:① 调整补全→Task1-3 ② 定期抓取方案→Task4 ✓ - 全集型简化避免过度设计(调查 agent 的回溯引擎是 future 时点型需求,现不做)✓ - TDD:migrate 聚合逻辑先写测试 ✓ - 不破坏:300/500/50/深证 migrate 路径不动,只加 000852/932000 announce 段 ✓ - 约束:不走代理/单线程/staging→merge 幂等/不碰 baostock ✓ ## 已知残留 gap(接受,不阻塞) - 932000 中间调整(2023-08 launch 到 current 之间)csindex 无公告 → launch ∪ current 全集,中间被踢的不可补(2023 新指数,影响小) - 000852 2014-2017 早期 HTML 表格解析格式松散,可能不全(扩 indexCode 搜索尽力补,实证 id=5/id=1585 等仍有表格)