feat(factor): P2 批 S 族情绪 4 因子 TDD(七项目调研增量收尾)——sentiment_lexicon 词表(分析师短语优先命中即定向+正/负词计数多数决,中文化移植 FinRobot 范式)+sentiment_adapter(news_meta 打分→次日可见 PIT eff=show_time+1 宁晚毋早;art_code 幂等去重;覆盖门 20 行窗新闻量<5→NaN 宁缺毋假;四因子 sent_mom_20/sent_chg[近30−前30 双窗各自门]/sent_burst[量 z 分,门窗均量≥0.5]/news_vol_chg[20/前20 量比])+sentiment_library 4 注册(category=sentiment 新类)+batch_eval 接线(sent_names 引用瘦身单帧 join,缺列全 NaN 降级);⚠️坑两记:①polars Boolean.sum() 产 UInt32,与 rolling_sum().over() 组合实测出现 −1 环回成 2^32−1 污染(值恰好多出 u32max+1)——计数列聚合即 cast Int64 消除;②测试 fixture 同日多行须按日分组再写 parquet(逐行覆盖写 part-00 会静默吃掉同日早行);前向积累型框定:corpus 2026-09-06 起采集,历史窗全 NaN 诚实透出,首个有意义 IC 窗=10 月下旬 Q3 披露潮;新增 13 用例,全套 1442 绿 [nas]
Co-Authored-By: Claude Code <noreply@anthropic.com>
This commit is contained in:
@@ -52,7 +52,7 @@
|
||||
| 内置 | ma/return 等基础 | 7 | 占位 |
|
||||
| 第一批 | vnpy Alpha101(**实挂 82**,18 个 add_feature 是注释行,grep 假数)+ Alpha158 | 240 注册/227 可算 | ✅ |
|
||||
| 财务批 | 聚宽财务 136 公式本地化(P0 32 个 TDD → P1 扩展) | 96 候选→32+37+6 | ✅ |
|
||||
| P2 批(09-22) | 七项目指标调研增量:D16 PEG/D17 PS + E14-16 股息族(dividend 域除息日锚:div_12m 365 天滚动/派息率/连续分红年数断年重计);PEG 分母走守卫列 peg_denom(np_ttm≤0 或 g≤1% → NaN 宁缺毋假);S 族情绪 4 因子随后批 | 5(+S 4 待) | ✅ TDD |
|
||||
| P2 批(09-22) | 七项目指标调研增量:D16 PEG/D17 PS + E14-16 股息族(dividend 域除息日锚:div_12m 365 天滚动/派息率/连续分红年数断年重计);PEG 分母走守卫列 peg_denom(np_ttm≤0 或 g≤1% → NaN 宁缺毋假);S 族情绪 4 因子(sentiment_lexicon 词表短语优先+计数多数决,news_meta 打分 PIT=次日可见,覆盖门窗内<5 条 NaN;**前向积累型**——corpus 2026-09-06 起采集,历史窗全 NaN 诚实透出,IC 随语料攒厚度,首个有意义窗=10 月下旬 Q3 披露潮) | 9 | ✅ TDD |
|
||||
| GTJA191 | — | — | ❄️ 冷冻(用户 09-12 拍板) |
|
||||
| Barra CNE5 | 风格因子 | — | 长线 |
|
||||
|
||||
|
||||
@@ -82,7 +82,9 @@ def run_batch_eval(
|
||||
# 分块进行——全量特征帧+alpha_df 双全量副本在 7.9G NAS 必 OOM)
|
||||
fund_names = [n for n in factor_names
|
||||
if (get_factor(n) or {}).get("category") in ("fundamental", "composite")]
|
||||
if fund_names:
|
||||
sent_names = [n for n in factor_names
|
||||
if (get_factor(n) or {}).get("category") == "sentiment"]
|
||||
if fund_names or sent_names:
|
||||
fund_static_dir = fund_data_dir or cfg.data_paths.get("static_dir") or DEFAULT_STATIC_DIR
|
||||
fund_codes = bars["vt_symbol"].unique().to_list()
|
||||
fund_days = bars["datetime"].unique().sort()
|
||||
@@ -169,6 +171,27 @@ def run_batch_eval(
|
||||
alpha_df = pl.concat(parts, how="vertical")
|
||||
parts.clear()
|
||||
|
||||
# 情绪因子批(S 族): corpus news_meta 单帧 join(新闻体量≪三表,不分块;
|
||||
# 引用瘦身同款——只 join 本批表达式引用的 SENTIMENT_COLUMNS 子集)
|
||||
if sent_names:
|
||||
import re as _re2
|
||||
from . import sentiment_library # noqa: F401 import 即注册
|
||||
from .sentiment_adapter import SENTIMENT_COLUMNS, build_sentiment_features
|
||||
needed_s = set()
|
||||
for n in sent_names:
|
||||
expr = (get_factor(n) or {}).get("expression", "")
|
||||
needed_s |= set(_re2.findall(r"[A-Za-z_][A-Za-z0-9_]*", expr)) \
|
||||
& set(SENTIMENT_COLUMNS)
|
||||
sent = build_sentiment_features(
|
||||
fund_codes, start, end, trading_dates=fund_days,
|
||||
columns=sorted(needed_s))
|
||||
if sent.height:
|
||||
alpha_df = alpha_df.join(sent, on=["vt_symbol", "datetime"], how="left")
|
||||
missing = [c for c in sorted(needed_s) if c not in alpha_df.columns]
|
||||
if missing:
|
||||
alpha_df = alpha_df.with_columns(
|
||||
[pl.lit(None, dtype=pl.Float64).alias(c) for c in missing])
|
||||
|
||||
for i, name in enumerate(factor_names):
|
||||
row = _eval_one(name, alpha_df, cutoff_map, start_dt, end_dt, rets,
|
||||
calculate_by_expression, values_out_dir=factor_values_out)
|
||||
|
||||
@@ -0,0 +1,182 @@
|
||||
# sanguo_factor/sentiment_adapter.py
|
||||
"""情绪因子适配层: corpus news_meta → PIT 日频特征列(S 族,2026-09-22 P2 批).
|
||||
|
||||
数据面: sanguo_data.corpus_reader.get_news_meta(title+summary 打分面;核内
|
||||
零去重→消费方按 art_code 幂等)。词表=sentiment_lexicon(短语优先+计数多数决).
|
||||
|
||||
口径红线:
|
||||
- PIT=次日可见: eff = show_time 所在日 + 1 天(当日盘后新闻泄入当日收盘评估的
|
||||
反例不可证,宁晚毋早——与重述可见性 notice=max 同族纪律);
|
||||
- 覆盖门(宁缺毋假,UZI 完整性门思想): 窗内新闻量 < MIN_NEWS → NaN 不硬算;
|
||||
- 前向积累型: corpus 采集 2026-09-06 起,历史窗无数据=全 NaN(诚实透出,
|
||||
IC 验证随语料攒厚度,首个有意义窗=10 月下旬 Q3 披露潮);
|
||||
- 窗口=网格行序(交易日网格≈月窗,与月度调仓对齐;日历网格=日历窗).
|
||||
|
||||
输出列(SENTIMENT_COLUMNS):
|
||||
- sent_mom_20 近 20 行窗净情绪 = Σ(pos−neg)/Σn(窗内全新闻聚合,门 n≥5)
|
||||
- sent_chg 近 30 行窗净情绪 − 前 30 行窗净情绪(两窗各自门 n≥5)
|
||||
- sent_burst 当日新闻量 z 分(对含当日的 20 行窗;门 窗均量≥0.5 且 std>0)
|
||||
- news_vol_chg 近 20 行窗量 / 前 20 行窗量 − 1(前窗 0 → NaN)
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import datetime, timedelta
|
||||
|
||||
import polars as pl
|
||||
|
||||
from .sentiment_lexicon import score_text
|
||||
|
||||
SENTIMENT_COLUMNS = ["sent_mom_20", "sent_chg", "sent_burst", "news_vol_chg"]
|
||||
|
||||
# 回看天数(日历): sent_chg 需前 30 行窗,交易日网格 ≈ 42 日历日,留边
|
||||
LOOKBACK_DAYS = 120
|
||||
MIN_NEWS = 5 # 覆盖门: 窗内新闻量下限
|
||||
|
||||
|
||||
def _vt_to_bare(vt: str) -> str:
|
||||
return vt.partition(".")[0]
|
||||
|
||||
|
||||
def _empty(columns: list[str]) -> pl.DataFrame:
|
||||
schema = {"vt_symbol": pl.Utf8, "datetime": pl.Datetime("us"),
|
||||
**{c: pl.Float64 for c in columns}}
|
||||
return pl.DataFrame(schema=schema)
|
||||
|
||||
|
||||
def _norm_day(v) -> str | None:
|
||||
"""show_time 任意形态(str/datetime/NaT)→ 'YYYY-MM-DD';无效 None."""
|
||||
if v is None:
|
||||
return None
|
||||
s = str(v)[:10]
|
||||
return s if len(s) == 10 and s[4] == "-" and s[:4].isdigit() else None
|
||||
|
||||
|
||||
def _series_or_blank(news, col: str):
|
||||
if col in news.columns:
|
||||
return news[col].fillna("")
|
||||
import pandas as pd
|
||||
return pd.Series([""] * len(news), index=news.index)
|
||||
|
||||
|
||||
def _scored_events(news, bare_to_vt: dict) -> pl.DataFrame:
|
||||
"""打分+次日可见 → (vt_symbol, eff, _s) 行集(无效行丢弃)."""
|
||||
titles = _series_or_blank(news, "title")
|
||||
summaries = _series_or_blank(news, "summary")
|
||||
vts, effs, ss = [], [], []
|
||||
for t, title, summ, code in zip(news["show_time"], titles, summaries,
|
||||
news["stock_code"]):
|
||||
day = _norm_day(t)
|
||||
vt = bare_to_vt.get(str(code))
|
||||
if day is None or vt is None:
|
||||
continue
|
||||
vts.append(vt)
|
||||
effs.append((datetime.strptime(day, "%Y-%m-%d")
|
||||
+ timedelta(days=1)).strftime("%Y-%m-%d"))
|
||||
ss.append(score_text(f"{title} {summ}"))
|
||||
if not vts:
|
||||
return pl.DataFrame(schema={"vt_symbol": pl.Utf8, "eff": pl.Utf8,
|
||||
"_s": pl.Int64})
|
||||
return pl.DataFrame({"vt_symbol": vts, "eff": effs, "_s": ss})
|
||||
|
||||
|
||||
def build_sentiment_features(
|
||||
codes: list[str],
|
||||
start: str,
|
||||
end: str,
|
||||
trading_dates: pl.Series | list | None = None,
|
||||
columns: list[str] | None = None,
|
||||
) -> pl.DataFrame:
|
||||
"""构建 PIT 日频情绪特征: vt_symbol × datetime × columns.
|
||||
|
||||
corpus 根=SANGUO_CORPUS_ROOT(corpus_reader 同源);缺数据/无新闻 → 空 schema
|
||||
帧不 raise(batch_eval join 后缺列=全 NaN,与财务域缺域降级同语义)。
|
||||
"""
|
||||
out_cols = list(columns) if columns is not None else list(SENTIMENT_COLUMNS)
|
||||
if not codes:
|
||||
return _empty(out_cols)
|
||||
try:
|
||||
from sanguo_data.corpus_reader import get_news_meta
|
||||
except Exception:
|
||||
return _empty(out_cols)
|
||||
lb_start = (datetime.strptime(start, "%Y-%m-%d")
|
||||
- timedelta(days=LOOKBACK_DAYS)).strftime("%Y-%m-%d")
|
||||
try:
|
||||
news = get_news_meta(start=lb_start, end=end,
|
||||
symbols=[_vt_to_bare(c) for c in codes])
|
||||
except Exception:
|
||||
return _empty(out_cols)
|
||||
if news is None or news.empty:
|
||||
return _empty(out_cols)
|
||||
news = news.drop_duplicates(subset=["art_code"])
|
||||
news = news[news["show_time"].notna() & news["stock_code"].notna()]
|
||||
if news.empty:
|
||||
return _empty(out_cols)
|
||||
|
||||
scored = _scored_events(news, {_vt_to_bare(c): c for c in codes})
|
||||
if scored.height == 0:
|
||||
return _empty(out_cols)
|
||||
daily = (scored.with_columns(
|
||||
pl.col("eff").str.to_date("%Y-%m-%d").alias("_d"))
|
||||
.group_by(["vt_symbol", "_d"]).agg(
|
||||
# cast Int64: Boolean.sum() 产 UInt32,rolling_sum().over() 组合下
|
||||
# 观测到 −1 环回成 2^32−1 的污染(值恰好多出 u32max+1,实测复现);
|
||||
# 计数量级极小,Int64 零成本消除整族 u32 滚动边界
|
||||
pl.len().cast(pl.Int64).alias("_n"),
|
||||
(pl.col("_s") > 0).sum().cast(pl.Int64).alias("_pos"),
|
||||
(pl.col("_s") < 0).sum().cast(pl.Int64).alias("_neg")))
|
||||
|
||||
# 日频 grid(codes × days);count 类缺日=0(事实性无新闻,非插补)
|
||||
if trading_dates is not None:
|
||||
days = sorted({str(d)[:10] for d in list(trading_dates)})
|
||||
else:
|
||||
s = datetime.strptime(start, "%Y-%m-%d")
|
||||
e = datetime.strptime(end, "%Y-%m-%d")
|
||||
days = []
|
||||
while s <= e:
|
||||
days.append(s.strftime("%Y-%m-%d"))
|
||||
s += timedelta(days=1)
|
||||
grid = pl.DataFrame({
|
||||
"vt_symbol": [c for c in codes for _ in days],
|
||||
"_d": days * len(codes),
|
||||
}).with_columns(pl.col("_d").str.to_date("%Y-%m-%d"))
|
||||
df = (grid.join(daily, on=["vt_symbol", "_d"], how="left")
|
||||
.with_columns(
|
||||
pl.col("_n").fill_null(0), pl.col("_pos").fill_null(0),
|
||||
pl.col("_neg").fill_null(0))
|
||||
.sort(["vt_symbol", "_d"]))
|
||||
|
||||
def _rsum(col: str, w: int, shift: int = 0) -> pl.Expr:
|
||||
e = pl.col(col).shift(shift) if shift else pl.col(col)
|
||||
return e.rolling_sum(w).over("vt_symbol")
|
||||
|
||||
n = pl.col("_n")
|
||||
df = df.with_columns(
|
||||
_rsum("_pos", 20).alias("_p20"), _rsum("_neg", 20).alias("_n20"),
|
||||
_rsum("_n", 20).alias("_c20"),
|
||||
_rsum("_pos", 30).alias("_p30"), _rsum("_neg", 30).alias("_n30"),
|
||||
_rsum("_n", 30).alias("_c30"),
|
||||
_rsum("_pos", 30, 30).alias("_p30p"), _rsum("_neg", 30, 30).alias("_n30p"),
|
||||
_rsum("_n", 30, 30).alias("_c30p"),
|
||||
_rsum("_n", 20, 20).alias("_c20p"),
|
||||
n.rolling_mean(20).over("vt_symbol").alias("_m20"),
|
||||
n.rolling_std(20).over("vt_symbol").alias("_sd20"),
|
||||
)
|
||||
net20 = pl.when(pl.col("_c20") >= MIN_NEWS).then(
|
||||
(pl.col("_p20") - pl.col("_n20")) / pl.col("_c20")).otherwise(None)
|
||||
net30 = pl.when(pl.col("_c30") >= MIN_NEWS).then(
|
||||
(pl.col("_p30") - pl.col("_n30")) / pl.col("_c30")).otherwise(None)
|
||||
net30p = pl.when(pl.col("_c30p") >= MIN_NEWS).then(
|
||||
(pl.col("_p30p") - pl.col("_n30p")) / pl.col("_c30p")).otherwise(None)
|
||||
df = df.with_columns(
|
||||
net20.alias("sent_mom_20"),
|
||||
(net30 - net30p).alias("sent_chg"),
|
||||
pl.when((pl.col("_m20") >= 0.5) & (pl.col("_sd20") > 0) & (n > 0))
|
||||
.then((n - pl.col("_m20")) / pl.col("_sd20")).otherwise(None)
|
||||
.alias("sent_burst"),
|
||||
pl.when(pl.col("_c20p") > 0)
|
||||
.then(pl.col("_c20") / pl.col("_c20p") - 1.0).otherwise(None)
|
||||
.alias("news_vol_chg"),
|
||||
)
|
||||
return (df.with_columns(pl.col("_d").cast(pl.Datetime("us")).alias("datetime"))
|
||||
.sort(["vt_symbol", "datetime"])
|
||||
.select(["vt_symbol", "datetime", *out_cols]))
|
||||
@@ -0,0 +1,64 @@
|
||||
# sanguo_factor/sentiment_lexicon.py
|
||||
"""情绪词表 v1(子串匹配零分词依赖,可解释可审计).
|
||||
|
||||
来源: 七项目指标调研(2026-09-22)——FinRobot catalyst_analyzer 英文词表
|
||||
(正 20/负 18)+分析师短语优先范式的中文化移植;阈值/口径属因子域场景语义
|
||||
(09-19 归属原则),策略域未来要消费再按 dsa 先例移交.
|
||||
|
||||
打分纪律(与 FinRobot 同构):
|
||||
- 分析师短语优先: 命中即定方向(短语比通用词信号强一个量级);
|
||||
- 否则正/负词命中**计数**多者胜,平局=0(单条新闻粒度);
|
||||
- 纯计数无权重——v1 接受噪声换可解释性,IC 见真章后再升级模型
|
||||
(升级路径: FinBERT/LLM,但数字纪律=词表或 LLM 均只产方向,数值另算).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
# 分析师/机构行动短语(优先级最高,命中即定方向)
|
||||
ANALYST_PATTERNS: dict[str, list[str]] = {
|
||||
"pos": [
|
||||
"上调评级", "上调至", "评级上调", "目标价上调", "上调目标价",
|
||||
"调高目标价", "首次覆盖", "予以买入", "维持买入", "维持增持",
|
||||
"看多", "强烈推荐",
|
||||
],
|
||||
"neg": [
|
||||
"下调评级", "下调至", "评级下调", "目标价下调", "下调目标价",
|
||||
"调低目标价", "维持卖出", "下调至持有", "维持减持", "看空",
|
||||
"不再覆盖", "剔除出",
|
||||
],
|
||||
}
|
||||
|
||||
# 通用正面词(计数制)
|
||||
POS_WORDS: tuple[str, ...] = (
|
||||
"超预期", "高于预期", "好于预期", "净利增长", "净利润增长", "业绩预增",
|
||||
"业绩增长", "营收增长", "中标", "签订合同", "签署合同", "订单",
|
||||
"回购", "增持", "分红", "创新高", "涨停", "涨价", "供不应求",
|
||||
"突破", "获批", "战略合作", "产能扩张", "扭亏", "盈利",
|
||||
)
|
||||
|
||||
# 通用负面词(计数制)
|
||||
NEG_WORDS: tuple[str, ...] = (
|
||||
"低于预期", "差于预期", "不及预期", "亏损", "预亏", "首亏", "业绩预减",
|
||||
"业绩下滑", "营收下滑", "负增长", "处罚", "立案", "调查", "违规",
|
||||
"减持", "质押爆仓", "平仓", "诉讼", "仲裁", "商誉减值", "减值",
|
||||
"退市", "警示", "违约", "债务逾期", "停产", "安全事故", "降价",
|
||||
"需求疲软", "下修",
|
||||
)
|
||||
|
||||
|
||||
def score_text(text: str) -> int:
|
||||
"""单条文本情绪分: +1/0/−1(短语优先,否则计数多数决,平局 0)."""
|
||||
if not text:
|
||||
return 0
|
||||
for p in ANALYST_PATTERNS["pos"]:
|
||||
if p in text:
|
||||
return 1
|
||||
for p in ANALYST_PATTERNS["neg"]:
|
||||
if p in text:
|
||||
return -1
|
||||
pos = sum(text.count(w) for w in POS_WORDS)
|
||||
neg = sum(text.count(w) for w in NEG_WORDS)
|
||||
if pos > neg:
|
||||
return 1
|
||||
if neg > pos:
|
||||
return -1
|
||||
return 0
|
||||
@@ -0,0 +1,31 @@
|
||||
# sanguo_factor/sentiment_library.py
|
||||
"""情绪因子批表达式库: S 族 4 个注册(category="sentiment",2026-09-22 P2 批).
|
||||
|
||||
选型来源: 七项目指标调研(2026-09-22)情绪增量层——FinRobot 词表范式/
|
||||
TradingAgents 量表语义/UZI 温度计思想的首批落地.原料=corpus news_meta
|
||||
(2026-09-06 起积累),前向积累型因子: 历史窗全 NaN 诚实透出,IC 验证
|
||||
随语料攒厚度(首个有意义窗=10 月下旬 Q3 披露潮).
|
||||
|
||||
方向: 全部按「高=好」正向注册,IC 待实证(词表打分的方向先验:
|
||||
情绪动量+ / 情绪变化+ / 爆发不定但事件窗有信息 / 关注度变化不定).
|
||||
"""
|
||||
from .registry import register_factor, _REGISTRY
|
||||
|
||||
# (name, expression, 调研编号, 方向先验)
|
||||
SENTIMENT_FACTORS: list[tuple[str, str, str, str]] = [
|
||||
("sent_mom_20", "cs_rank(sent_mom_20)", "S01", "+"),
|
||||
("sent_chg", "cs_rank(sent_chg)", "S02", "+"),
|
||||
("sent_burst", "cs_rank(sent_burst)", "S03", "不定"),
|
||||
("news_vol_chg", "cs_rank(news_vol_chg)", "S04", "不定"),
|
||||
]
|
||||
|
||||
|
||||
def _register_all() -> None:
|
||||
"""注册全部情绪因子(已存在同名跳过,幂等;同 fundamental_library 模式)."""
|
||||
for name, expression, _doc_id, _ic in SENTIMENT_FACTORS:
|
||||
if name not in _REGISTRY:
|
||||
register_factor(name, expression, category="sentiment")
|
||||
|
||||
|
||||
# 模块导入时自动注册(与 library.py/fundamental_library.py 同一模式)
|
||||
_register_all()
|
||||
@@ -0,0 +1,169 @@
|
||||
# tests/factor/test_sentiment_adapter.py
|
||||
"""S 族情绪因子适配层: 词表打分 + corpus news_meta → PIT 日频特征.
|
||||
|
||||
口径锚(P2 任务书=七项目调研报告 §5.2.3,2026-09-22):
|
||||
- 词表: 分析师短语优先(命中即定方向),否则正/负词计数多数决,平局 0
|
||||
- PIT=次日可见: eff = show_time 所在日 + 1(宁晚毋早)
|
||||
- art_code 幂等去重(核内零去重契约)
|
||||
- 覆盖门: 20 行窗内新闻量 <5 → NaN(宁缺毋假,不硬算)
|
||||
- 前向积累型: 无新闻股全列 NaN(≠0)
|
||||
"""
|
||||
import math
|
||||
import sys, os
|
||||
sys.path.insert(0, os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..")))
|
||||
sys.path.insert(0, os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..", "vnpy_v4.4.0")))
|
||||
|
||||
import pandas as pd
|
||||
import polars as pl
|
||||
import pytest
|
||||
|
||||
from sanguo_factor.sentiment_lexicon import score_text
|
||||
from sanguo_factor.sentiment_adapter import build_sentiment_features
|
||||
|
||||
A, B, C = "600000.SSE", "000001.SZSE", "300001.SZSE"
|
||||
|
||||
|
||||
# ---------- 词表单元 ----------
|
||||
|
||||
def test_phrase_priority_over_words():
|
||||
# 短语优先: 「上调评级」+多个负面词 → 仍 +1
|
||||
assert score_text("机构上调评级,尽管此前有处罚与诉讼纠纷") == 1
|
||||
assert score_text("下调评级,虽然公司中标新项目") == -1
|
||||
|
||||
|
||||
def test_count_majority_and_tie():
|
||||
assert score_text("公司中标重大项目") == 1
|
||||
assert score_text("业绩下滑且收到处罚") == -1
|
||||
assert score_text("召开股东大会审议年度报告") == 0 # 无命中
|
||||
assert score_text("回购遇上减持") == 0 # 1:1 平局
|
||||
|
||||
|
||||
# ---------- 合成 corpus 树 ----------
|
||||
|
||||
def _news(art, code, show, title):
|
||||
return {"art_code": art, "stock_code": code,
|
||||
"show_time": f"{show} 10:30:00", "title": title, "summary": "",
|
||||
"media_name": "测试社", "url": "", "first_seen_date": show}
|
||||
|
||||
|
||||
_A_NEWS = [
|
||||
# 12 月六连(A06 eff=12-24=PIT 边界用例: 12-23 不可见/12-24 可见)
|
||||
_news("A01", "600000", "2025-12-04", "公司中标重大项目"),
|
||||
_news("A02", "600000", "2025-12-08", "营收增长30%"),
|
||||
_news("A03", "600000", "2025-12-12", "被立案调查"),
|
||||
_news("A04", "600000", "2025-12-16", "公司签订合同"),
|
||||
_news("A05", "600000", "2025-12-20", "业绩下滑"),
|
||||
_news("A06", "600000", "2025-12-23", "年报扭亏"),
|
||||
# 1 月九连 + 爬虫跨天重复行(同 art_code,须幂等)
|
||||
_news("A07", "600000", "2026-01-04", "回购股份方案公布"),
|
||||
_news("A08", "600000", "2026-01-07", "净利润增长20%"),
|
||||
_news("A09", "600000", "2026-01-11", "收到监管处罚"),
|
||||
_news("A10", "600000", "2026-01-14", "多家机构上调评级"),
|
||||
_news("A11", "600000", "2026-01-19", "大股东减持公告"),
|
||||
_news("A11", "600000", "2026-01-20", "大股东减持公告"), # dup(同 art_code)
|
||||
_news("A12", "600000", "2026-01-21", "召开股东大会"),
|
||||
_news("A13", "600000", "2026-01-24", "计提商誉减值"),
|
||||
_news("A14", "600000", "2026-01-26", "四季度扭亏"),
|
||||
_news("A15", "600000", "2026-01-28", "收到警示函"),
|
||||
]
|
||||
# B: 12 连 neutral + 01-20 爆发日 4 条(全 neutral → mom=0,burst 纯量效应)
|
||||
_B_NEWS = ([_news(f"B{i:02d}", "000001", f"2026-01-{d:02d}", "召开股东大会")
|
||||
for i, d in enumerate(range(1, 13), 1)]
|
||||
+ [_news(f"BX{i}", "000001", "2026-01-19", "公司发布公告") for i in range(4)])
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def corpus_root(tmp_path, monkeypatch):
|
||||
root = tmp_path / "corpus"
|
||||
dom = root / "news_meta"
|
||||
# 同日多行必须先分组再写(逐行覆盖写 part-00.parquet 会把同日早行吃掉——
|
||||
# B 爆发日 4 条/A 与 B 共享分区日均中过招)
|
||||
by_day: dict[str, list] = {}
|
||||
for row in _A_NEWS + _B_NEWS:
|
||||
by_day.setdefault(row["show_time"][:10], []).append(row)
|
||||
for day, rows in by_day.items():
|
||||
pdir = dom / f"dt={day}"
|
||||
pdir.mkdir(parents=True, exist_ok=True)
|
||||
pd.DataFrame(rows).to_parquet(pdir / "part-00.parquet", index=False)
|
||||
monkeypatch.setenv("SANGUO_CORPUS_ROOT", str(root))
|
||||
return root
|
||||
|
||||
|
||||
def _val(df, vt, day, col):
|
||||
import datetime as _d
|
||||
row = df.filter((df["vt_symbol"] == vt)
|
||||
& (df["datetime"] == _d.datetime.strptime(day, "%Y-%m-%d")))
|
||||
assert row.height == 1, f"grid 缺行 {vt} {day}"
|
||||
v = row[col][0]
|
||||
return None if v is None else float(v)
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def feat_build():
|
||||
return build_sentiment_features
|
||||
|
||||
|
||||
def test_pit_next_day_visibility(corpus_root):
|
||||
df = build_sentiment_features([A], "2025-11-01", "2026-01-31")
|
||||
# A06(show 12-23)eff=12-24: 12-23 窗(12-04..12-23)5 条 pos3 neg2 → 0.2;
|
||||
# 12-24 窗含它(6 条 pos4 neg2 → 1/3)——次日可见边界(当日盘后新闻不泄入当日)
|
||||
assert _val(df, A, "2025-12-23", "sent_mom_20") == pytest.approx(0.2)
|
||||
assert _val(df, A, "2025-12-24", "sent_mom_20") == pytest.approx(1.0 / 3.0)
|
||||
|
||||
|
||||
def test_mom_window_value_and_gate(corpus_root):
|
||||
df = build_sentiment_features([A], "2025-11-01", "2026-01-31")
|
||||
# 01-31 的 20 行窗(01-12..01-31): eff 7 条(pos2 neg4;A11 重复行已去重)
|
||||
assert _val(df, A, "2026-01-31", "sent_mom_20") == pytest.approx(-2.0 / 7.0)
|
||||
# 覆盖门: 01-10 窗(12-22..01-10)仅 3 条 → NaN
|
||||
assert _val(df, A, "2026-01-10", "sent_mom_20") is None
|
||||
|
||||
|
||||
def test_art_code_dedup_in_window(corpus_root):
|
||||
df = build_sentiment_features([A], "2025-11-01", "2026-01-31")
|
||||
# 若 dup 未去重: 01-21 多一条 neg → mom=(2-5)/8;去重后 -2/7
|
||||
assert _val(df, A, "2026-01-31", "sent_mom_20") == pytest.approx(-2.0 / 7.0)
|
||||
|
||||
|
||||
def test_sent_chg_two_window_difference(corpus_root):
|
||||
df = build_sentiment_features([A], "2025-11-01", "2026-01-31")
|
||||
# 01-31: net30(01-02..01-31)=9 条 pos4 neg4 → 0;net30p(12-03..01-01)
|
||||
# =6 条 pos4 neg2 → 1/3;chg = −1/3
|
||||
assert _val(df, A, "2026-01-31", "sent_chg") == pytest.approx(-1.0 / 3.0)
|
||||
|
||||
|
||||
def test_news_vol_chg_ratio(corpus_root):
|
||||
df = build_sentiment_features([A], "2025-11-01", "2026-01-31")
|
||||
# 01-31: c20=7(01-12..01-31),c20p=3(12-23..01-11 窗: 12-24/01-05/01-08)
|
||||
# → 7/3−1 = 4/3
|
||||
assert _val(df, A, "2026-01-31", "news_vol_chg") == pytest.approx(4.0 / 3.0)
|
||||
|
||||
|
||||
def test_burst_zscore_and_gates(corpus_root):
|
||||
df = build_sentiment_features([A, B], "2025-11-01", "2026-01-31")
|
||||
# B 01-20(eff): 窗 20 行 = 12×1 + 4 + 7×0 → mean=0.8, var(ddof=1)=0.8
|
||||
# z = (4−0.8)/sqrt(0.8)
|
||||
assert _val(df, B, "2026-01-20", "sent_burst") == pytest.approx(
|
||||
3.2 / math.sqrt(0.8))
|
||||
# 非新闻日 → NaN(门 n>0)
|
||||
assert _val(df, B, "2026-01-15", "sent_burst") is None
|
||||
# 窗均量不足(0.5 门)→ NaN
|
||||
assert _val(df, A, "2026-01-31", "sent_burst") is None
|
||||
|
||||
|
||||
def test_no_news_stock_all_nan(corpus_root):
|
||||
df = build_sentiment_features([A, C], "2025-11-01", "2026-01-31")
|
||||
for col in ("sent_mom_20", "sent_chg", "sent_burst", "news_vol_chg"):
|
||||
assert _val(df, C, "2026-01-15", col) is None, f"{col} 无新闻股应 NaN 非 0"
|
||||
|
||||
|
||||
def test_columns_subset(corpus_root):
|
||||
df = build_sentiment_features([A], "2026-01-01", "2026-01-31",
|
||||
columns=["sent_mom_20"])
|
||||
assert df.columns == ["vt_symbol", "datetime", "sent_mom_20"]
|
||||
|
||||
|
||||
def test_empty_corpus_returns_schema_frame(tmp_path, monkeypatch):
|
||||
monkeypatch.setenv("SANGUO_CORPUS_ROOT", str(tmp_path / "nothing"))
|
||||
df = build_sentiment_features([A], "2026-01-01", "2026-01-31")
|
||||
assert df.height == 0 and "sent_mom_20" in df.columns
|
||||
@@ -0,0 +1,36 @@
|
||||
# tests/factor/test_sentiment_library.py
|
||||
"""S 族情绪因子表达式库: 4 个注册与契约锁定(P2 批)."""
|
||||
import re
|
||||
import sys, os
|
||||
sys.path.insert(0, os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..")))
|
||||
sys.path.insert(0, os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..", "vnpy_v4.4.0")))
|
||||
|
||||
import pytest
|
||||
|
||||
from sanguo_factor import sentiment_library # noqa: F401 import 即注册
|
||||
from sanguo_factor.sentiment_adapter import SENTIMENT_COLUMNS
|
||||
from sanguo_factor.sentiment_library import SENTIMENT_FACTORS
|
||||
from sanguo_factor.registry import list_factors, get_factor
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _ensure_registered():
|
||||
sentiment_library._register_all()
|
||||
|
||||
|
||||
def test_s4_factors_registered():
|
||||
facs = {f["name"]: f for f in list_factors("sentiment")}
|
||||
names = {n for n, _e, _d, _ic in SENTIMENT_FACTORS}
|
||||
assert names == {"sent_mom_20", "sent_chg", "sent_burst", "news_vol_chg"}
|
||||
for n in names:
|
||||
assert n in facs and facs[n]["category"] == "sentiment"
|
||||
assert facs[n]["expression"].startswith("cs_rank(")
|
||||
assert facs[n]["expression"].count("cs_rank(") == 1
|
||||
|
||||
|
||||
def test_expressions_reference_only_sentiment_columns():
|
||||
allowed = set(SENTIMENT_COLUMNS) | {"cs_rank"}
|
||||
for name, expr, _doc, _ic in SENTIMENT_FACTORS:
|
||||
bare = re.findall(r"[A-Za-z_][A-Za-z0-9_]*", expr)
|
||||
unknown = [t for t in bare if t not in allowed]
|
||||
assert not unknown, f"{name} 引用未知列 {unknown}"
|
||||
Reference in New Issue
Block a user