Files

44 lines
2.0 KiB
Python
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# -*- coding: utf-8 -*-
"""dmsk 财务五域薄封装(spec §20.11.3)。
schema 契约(2026-09-17 NAS 实测落档,§20.11.3 数据契约):
- 五域每行 NOTICE_DATE 必在(实测全命中;缺失=写路径事故,消费方按整期
弃用=宁缺毋假红线)
- 代码双列:SECUCODE 带后缀(301569.SZ)/ SECURITY_CODE 裸 6 位
- 金额单位=元(EM datacenter 原生,非万元)
- REPORT_DATE=分区键同值(dt=<REPORT_DATE>)
noticed_by=纯 filter 非 PIT merge(文档明示):只做 NOTICE_DATE<=noticed_by
的行级过滤(日期前缀比较);「static 基线为主+dmsk 新鲜度补丁」的 PIT 合并
归因子侧 fundamental_pit(§19.3 下游读法),不在数据线。
noticed_by 语义备注(因子线 09-18 review 记录②):NOTICE_DATE 为空/NaT 的行
经前缀比较会被滤除('NaT' 字典序>任何日期串)——**有意语义**=宁缺毋假
(无法证明披露可见即不可见),非 bug 勿修。
"""
import sanguo_data.partition_reader as pr
DOMAINS = ("balance", "cashflow", "express", "forecast", "income")
def get_domain(domain, start=None, end=None, noticed_by=None):
"""读五域之一(dt=REPORT_DATE 分区直读,零转换零去重)。
start/end 过滤的是 REPORT_DATE 分区值;noticed_by 过滤 NOTICE_DATE
(接受 'YYYY-MM-DD' 或全时间戳,按日期前缀比较)。
"""
if domain not in DOMAINS:
raise ValueError("domain 必须是 %s 之一, 收到 %r" % (DOMAINS, domain))
df = pr.read_partitions(pr.fund_dmsk_root(), domain, start, end)
if noticed_by is not None and not df.empty:
if "NOTICE_DATE" not in df.columns:
raise KeyError(
"域 %s 缺 NOTICE_DATE 列(数据契约①违约, 整期弃用勿消费)" % domain)
df = df[df["NOTICE_DATE"].astype(str).str[:10] <= noticed_by[:10]]
return df
def get_watermark():
"""镜像水位线(None=尚无成功推送记录)。"""
return pr.read_watermark(pr.fund_dmsk_root())