Commit b783f6ff authored by Data Governance Dev's avatar Data Governance Dev

feat(standards): 公积金领域 8 项新 indicator + 字段注释匹配

字段匹配双路化(comment_keywords 优先于 applies_to_fields):
  - standards/base.py: BaseStandard 加 comment_keywords: list[str]
  - web/core/step_impl/step7_standards.py: 新增 _collect_hit_columns()
    第 1 路按字段名命中,第 2 路按字段注释 substring 命中(不区分大小写)
    两路并集去重(按 (table, column))
  - standards/registry.py: list_all_standards() 输出 comment_keywords

8 个 indicator 新增/修改:
  - IND-003-a (改): 加「不可全 0」前置校验
  - IND-005-a (新): GB/T 15835 固定电话(区号 3~4 位 + 号码 7~8 位)
  - IND-006-a (新): 通讯地址(长度 [4,200],非全空白/符号/数字)
  - IND-006-b (新): GB/T 23705 邮编(6 位数字)
  - IND-007-a (新): 个人公积金账号(仅位数 check,10/12/18 位,applies_to_fields 故意留空)
  - IND-008-a (新): 首次参加工作年月(YYYYMM/MMDD)
  - IND-009-a (新): 用工类型(6 值枚举 + 文字形式)
  - IND-010-a (新): 本地户籍标记(1/0、是/否、Y/N、T/F、TRUE/FALSE)
  - IND-011-a (新): 职工状态(5 值枚举 + 流转规则说明)

orchestrator / 前端:
  - web/core/orchestrator.py: _derive_target / _derive_check 加 8 个新 indicator 分支;
    _register_indicator_steps() description 同时显示字段名 + 注释关键字
  - web/configs/analysis_tree.json: 注册 8 个新 step_id;新增「业务字段规范」组(默认展开)
  - 国标字段规范组加 std_ind_005_a

KPI 去 card 化:
  - web/static/style.css: 去掉 .kpi-card 的 background / shadow / border / radius / padding;
    改为数字带色 + label 灰字的纯文字布局,密排紧凑

已知限制:
  - IND-011-a 流转限制(销户 → 不可缴存、封存 → 可偿贷)需跨表 SQL JOIN,
    本轮仅在 description 写明业务规则,留待后续跨表 indicator

WORKLOG: docs/WORKLOG.md 2026-08-10 第二条
parent 1c07dd8c
......@@ -2,6 +2,100 @@
> 任务做完一次记一次。最近的在最上面。
## 2026-08-10 · 公积金领域 8 项新 indicator + 字段注释匹配
### 用户原话
> 对于字段匹配都先用字段注释匹配
> 2.1 修改(手机号加「不可全0」)
> 2.2 实现(固定电话)
> 2.3 实现(通讯地址+邮编)
> 2.4 仅实限位数check,用字段注释匹配(个人公积金账号)
> 3.1 实现(首次参加工作年月)
> 3.2 实现(用工类型)
> 3.3 实现(本地户籍标记)
> 3.4 实现(职工状态 + 流转)
### 关键架构决策:comment_keywords 优先于 applies_to_fields
业务字段名五花八门(mobile / phone / tel / lxdh / gjjzh ...),但字段注释("手机号码"/"联系电话"/"个人公积金账号")往往更稳定 —— 用户要求「先用字段注释匹配」。
实现:BaseStandard 新增 `comment_keywords: list[str]`,与 `applies_to_fields` 并集命中;任一非空即可生效。
### 改动总览
#### 1. BaseStandard 加 comment_keywords 字段
- [standards/base.py:55-58](standards/base.py#L55-L58) `BaseStandard.comment_keywords: list[str] | None`
- `__init__` 兜底 None → []
#### 2. step7 字段匹配双路化
- [web/core/step_impl/step7_standards.py:484-516](web/core/step_impl/step7_standards.py#L484-L516) 新增 `_collect_hit_columns()`:
- 第 1 路:按字段名匹配 (`applies_to_fields`)
- 第 2 路:按字段注释 substring 匹配(`comment_keywords`,不区分大小写)
- 两路并集去重(按 `(table, column)`)
- [web/core/step_impl/step7_standards.py:527-564](web/core/step_impl/step7_standards.py#L527-L564) `_run_round1_one()` 改签名:by_name → all_columns(接收完整字段元数据)
- 单 (table, field) 命中后保留 MAX_TABLES_PER_FIELD 截断
#### 3. 8 个 indicator 新增/修改
| ID | 名称 | 字段匹配方式 | 备注 |
|----|------|--------------|------|
| IND-003-a | 工信部 手机号格式(修改) | 字段名 + 注释 | 加「不可全0」前置检查 |
| IND-005-a | GB/T 15835 固定电话格式 | 字段名 + 注释 | 区号 3~4 位 + 号码 7~8 位 |
| IND-006-a | 通讯地址 格式 | 字段名 + 注释 | 长度 4~200 + 非全空白/全符号/全数字 |
| IND-006-b | GB/T 23705 邮政编码 格式 | 字段名 + 注释 | 6 位数字 |
| IND-007-a | 个人公积金账号 格式(位数) | **仅注释** | 故意留空 applies_to_fields |
| IND-008-a | 首次参加工作年月 格式 | 字段名 + 注释 | 6/8 位 YYYYMM(MMDD) |
| IND-009-a | 用工类型 值域 | 字段名 + 注释 | 6 值枚举 + 文字形式 |
| IND-010-a | 本地户籍标记 值域 | 字段名 + 注释 | 1/0/是/否/Y/N/T/F |
| IND-011-a | 职工状态 值域 | 字段名 + 注释 | 5 值枚举 + 流转说明 |
#### 4. orchestrator 适配
- [web/core/orchestrator.py:271-287](web/core/orchestrator.py#L271-L287) `_derive_target()` 加 8 个新 indicator 的特殊分支;comment-only 类显式标记「字段注释含...」
- [web/core/orchestrator.py:326-329](web/core/orchestrator.py#L326-L329) `_derive_check()` special_checks 表加 8 个新分支(多行 / 含 GB 号 / 含算法摘要)
- `_register_indicator_steps()` description 同时显示 `applies_to_fields` + `comment_keywords`
#### 5. 前端注册
- [web/configs/analysis_tree.json:33-37](web/configs/analysis_tree.json#L33-L37) 「国标字段规范」加 std_ind_005_a
- [web/configs/analysis_tree.json:67-78](web/configs/analysis_tree.json#L67-L78) 新增「业务字段规范」组(含 IND-006/007/008/009/010/011 系列)
- 默认展开,让用户第一时间看到新检查
#### 6. registry
- [standards/registry.py:78](standards/registry.py#L78) `list_all_standards()` 输出 `comment_keywords`,前端 / 报告可消费
### 已知限制
- **IND-011-a 流转限制**(销户 → 不可缴存、封存 → 不可贷款)需跨表 SQL JOIN 校验,本 indicator 仅在 description 写明规则 —— 留给后续跨表 indicator(参考 IND-301/302 模板)
- IND-007-a 的字段匹配**故意**不走字段名(不同业务系统命名差异大),仅走注释匹配 —— 测试已确认(`staff.gjjzh` 注释"个人公积金账号"成功命中)
### 测试覆盖
- 8 个 indicator 的 `validate()` 各跑了 10~14 个用例,全部预期命中
- `_collect_hit_columns()` 双路命中验证:
- IND-005-a 命中 3 个字段(字段名 1 + 注释 2)
- IND-007-a 仅命中 1 个字段(注释路径)
- IND-009-a 命中 1 个字段(字段名)
- 36 个 indicator 自动发现(32 IND + 4 STD 已 deprecated)
### 文件清单
新建 7 个:
- [standards/ind_005a_fixed_phone_format.py](standards/ind_005a_fixed_phone_format.py)
- [standards/ind_006a_address_format.py](standards/ind_006a_address_format.py)
- [standards/ind_006b_postal_code_format.py](standards/ind_006b_postal_code_format.py)
- [standards/ind_007a_fund_account_format.py](standards/ind_007a_fund_account_format.py)
- [standards/ind_008a_first_work_month_format.py](standards/ind_008a_first_work_month_format.py)
- [standards/ind_009a_employment_type_format.py](standards/ind_009a_employment_type_format.py)
- [standards/ind_010a_local_household_flag.py](standards/ind_010a_local_household_flag.py)
- [standards/ind_011a_employee_status_format.py](standards/ind_011a_employee_status_format.py)
修改 6 个:
- [standards/base.py](standards/base.py) 加 comment_keywords
- [standards/ind_003a_mobile_format.py](standards/ind_003a_mobile_format.py) 加全 0 检查 + 注释匹配
- [standards/registry.py](standards/registry.py) list_all 输出 comment_keywords
- [web/core/step_impl/step7_standards.py](web/core/step_impl/step7_standards.py) 双路命中
- [web/core/orchestrator.py](web/core/orchestrator.py) 8 indicator 分支 + _register_indicator_steps description
- [web/configs/analysis_tree.json](web/configs/analysis_tree.json) 注册 8 个 step_id
---
## 2026-08-10 · 检查项 info icon:扩 4 段 + 修两个 root cause
### 用户原话(这一轮两个迭代)
......
......@@ -32,10 +32,17 @@ class BaseStandard(ABC):
子类需实现:
- standard_id: 标准编号,如 "IND-001-a"
- standard_name: 标准中文名
- applies_to_fields: 适用的字段名(list)
- applies_to_fields / comment_keywords: 至少有一个非空,详见下
- validate(value): 单值校验逻辑
- describe(): 给人看的描述(可选)
字段匹配策略(指示字段 —— 决定「跑哪些字段」):
- applies_to_fields: 按字段名匹配(list[str],精确字符串)
- comment_keywords: 按字段注释 substring 匹配(list[str],不区分大小写)
两者并集(去重)作为命中集合。任何一个非空即可生效。
实际业务中字段名五花八门(mobile / phone / tel / lxdh ...),
字段注释("手机号码"/"联系电话")往往更稳定 —— 故推荐以注释匹配为主。
扩展元数据(可选覆盖):
- group: "国标字段规范" / "证件信息" / "身份信息" / "民政信息" / "通用"
- severity: "warning" | "error" —— 影响前端 KPI 染色
......@@ -49,6 +56,8 @@ class BaseStandard(ABC):
standard_id: str = ""
standard_name: str = ""
applies_to_fields: list[str] | None = None
# 字段注释关键字(中文 substring 包含即命中,不区分大小写)
comment_keywords: list[str] | None = None
description: str = ""
# ── 扩展元数据(向后兼容 —— 缺省走「通用 / warning」) ──
......@@ -65,6 +74,8 @@ class BaseStandard(ABC):
# 子类可省略 __init__;基类负责把 None 兜底为 []
if self.applies_to_fields is None:
self.applies_to_fields = []
if self.comment_keywords is None:
self.comment_keywords = []
if self.depends_on is None:
self.depends_on = []
......
"""IND-003-a 手机号格式(11 位 + 1[3-9] 开头 + 清洗)
"""IND-003-a 手机号格式(11 位 + 1[3-9] 开头 + 清洗 + 不可全 0)
拆分自 STD-003。仅做宽松长度 / 头部校验。
拆分自 STD-003。仅做宽松长度 / 头部 / 异常值校验。
校验要点:
1. 自动清洗:去 +86 / 86 前缀、去空格 / 横线
2. 长度清洗后 = 11
3. 头部 1[3-9](第二位 3-9)
4. **不可全 0**(如 "00000000000"、"12345600000" 里的号段残缺值)—— 业务上视为无效
"""
from __future__ import annotations
......@@ -18,7 +24,11 @@ class MobileFormatIndicator(BaseStandard):
"contact_phone", "customer_mobile", "legal_person_mobile",
"notify_phone", "maintainer_phone", "manager_phone", "receiver_phone",
]
description = "11 位;以 1 开头,第 2 位 3-9;自动清洗 +86 / 空格 / 横线"
comment_keywords = ["手机号码", "手机号", "联系手机", "联系电话"]
description = (
"11 位;以 1 开头,第 2 位 3-9;自动清洗 +86 / 空格 / 横线;"
"**不可全 0**(如 00000000000 / 00000123456 等号段残缺值视为无效)"
)
group = "国标字段规范"
severity = "warning"
......@@ -39,6 +49,10 @@ class MobileFormatIndicator(BaseStandard):
v = self._clean(value)
if len(v) != 11:
return ValidationResult(False, "", value, f"清洗后长度 {len(v)} ≠ 11", self.standard_id)
# 「不可全 0」前置拦截:00000000000 / 00000123456 等号段残缺值,
# 业务上视为无效 —— 给更具体的 reason,方便定位数据源问题
if set(v) == {"0"}:
return ValidationResult(False, "", value, "全 0 值(00000000000),业务上无效", self.standard_id)
if not self._REGEX.match(v):
return ValidationResult(False, "", value, "不符合 1[3-9]XXXXXXXXX 格式", self.standard_id)
return ValidationResult(True, "", value, "", self.standard_id)
\ No newline at end of file
"""IND-005-a 固定电话格式(区号 + 号码)
参考:
- GB/T 15835《电话号码编码规则》未公开全文,业内一般按以下口径校验:
1. 区号:以 0 开头,长度 3~4 位(如 010、021、0755、0835)
2. 号码:7~8 位
3. 区号与号码间允许 0~1 个分隔符(短横线 - 或空格)
不做:
- 号码段存在性(工信部未公开号段清单,仅手机号有,固话无)
- 国际长途(0086 / +86 等跳过 —— 业务中固定电话一般为国内号)
样例:
✓ 010-12345678 021-1234567 0755-12345678 01012345678
✗ 110 12345 010-123 010-12abc3456
"""
from __future__ import annotations
import re
from .base import BaseStandard, ValidationResult
class FixedPhoneFormatIndicator(BaseStandard):
standard_id = "IND-005-a"
standard_name = "GB/T 15835 固定电话格式"
applies_to_fields = [
"fixed_phone", "office_phone", "tel", "fax",
"home_phone", "work_phone", "company_phone", "phone_office",
]
comment_keywords = [
"固定电话", "办公电话", "公司电话", "单位电话", "工作电话",
"联系电话", "传真", "座机",
]
description = (
"0 开头区号(3~4 位)+ 可选「- / 空格」分隔 + 7~8 位号码;"
"仅校验长度 + 头部,不查号段(工信部未公开固话号段表)"
)
group = "国标字段规范"
severity = "warning"
# 区号 3~4 位(必 0 开头)+ 可选 - / 空格 + 号码 7~8 位
_REGEX = re.compile(r"^0\d{2,3}[-\s]?\d{7,8}$")
@staticmethod
def _clean(raw: str) -> str:
# 去掉括号(部分系统录入 "(010)12345678")和明显空格
v = raw.replace("(", "").replace(")", "")
return v.strip()
def validate(self, value: str) -> ValidationResult:
if not value:
return ValidationResult(True, "", value, "空值跳过", self.standard_id)
v = self._clean(value)
if not self._REGEX.match(v):
return ValidationResult(
False, "", value,
"不符合「0XX/0XXX - 7~8 位」固话格式",
self.standard_id,
)
return ValidationResult(True, "", value, "", self.standard_id)
\ No newline at end of file
"""IND-006-a 通讯地址格式
不做严格的「行政区划 / 路名 / 门牌号」语义校验(无公开标准),只做:
1. 非空(trim 后长度 ≥ 4 —— 至少含「省/市/区」3 字 + 1 字其它)
2. 不可全空白 / 全符号 / 全数字
3. 长度上限 200 字符(多数业务系统 VARCHAR(200),超出即被截断)
4. 不含常见录入错误字符(全角空格、转义符等 —— 极端兜底)
参考:
- GB/T 23705-2009《地名 邮政编码》未对地址文本本身定硬性规则
- 业内普遍用「省/市/区/路/号」关键字判定 —— 本 indicator 不走 LLM,纯规则
"""
from __future__ import annotations
from .base import BaseStandard, ValidationResult
class AddressFormatIndicator(BaseStandard):
standard_id = "IND-006-a"
standard_name = "通讯地址 格式"
applies_to_fields = [
"address", "mailing_address", "home_address", "contact_address",
"register_address", "company_address", "work_address", "live_address",
"addr", "postal_address",
]
comment_keywords = [
"通讯地址", "联系地址", "户籍地址", "居住地址", "住址", "现住址",
"工作地址", "单位地址", "公司地址", "注册地址", "办公地址",
"通讯地址", "收件地址", "邮寄地址", "送达地址", "通讯地点",
]
description = (
"trim 后 4 ≤ len ≤ 200;不全空白 / 不全符号 / 不全数字;"
"纯规则(无国标对应,不查行政区划存在性)"
)
group = "通用"
severity = "warning"
_MIN_LEN = 4
_MAX_LEN = 200
def validate(self, value: str) -> ValidationResult:
if value is None or str(value) == "":
return ValidationResult(True, "", value, "空值跳过", self.standard_id)
v = str(value).strip()
if not v:
return ValidationResult(False, "", value, "全空白", self.standard_id)
if len(v) < self._MIN_LEN:
return ValidationResult(
False, "", value,
f"trim 后长度 {len(v)} < {self._MIN_LEN}(疑似行政区划不全)",
self.standard_id,
)
if len(v) > self._MAX_LEN:
return ValidationResult(
False, "", value,
f"长度 {len(v)} > {self._MAX_LEN}(疑似录入异常 / 字段定义过短)",
self.standard_id,
)
# 全数字 / 全符号 / 全字母 → 视为异常地址
# 「字母」包含汉字外的字符;汉字在 Unicode >= 0x4E00
has_chinese = any("一" <= ch <= "鿿" for ch in v)
has_digit = any(ch.isdigit() for ch in v)
# 字母仅算 A-Z/a-z;汉字不算 letter
has_alpha = any(("a" <= ch <= "z") or ("A" <= ch <= "Z") for ch in v)
# 标点 / 空格单独算「有内容」
non_alnum_chinese = [ch for ch in v if not (
"一" <= ch <= "鿿"
or ch.isdigit()
or ("a" <= ch <= "z") or ("A" <= ch <= "Z")
)]
if not (has_chinese or has_digit or has_alpha or non_alnum_chinese):
return ValidationResult(False, "", value, "无可识别字符", self.standard_id)
# 全数字(且无汉字)→ 异常(如 "12345")
if not has_chinese and not has_alpha and has_digit:
return ValidationResult(False, "", value, "全数字且无汉字(疑似区号 / 编号)", self.standard_id)
# 含典型占位 / 录入错误关键字
bad_markers = ("无", "暂无", "null", "NULL", "未知", "TBD", "tbd", "/")
if v in bad_markers or all(m in v for m in ("无",)):
return ValidationResult(False, "", value, "录入占位符(无 / 暂无 / 未知 / TBD)", self.standard_id)
return ValidationResult(True, "", value, "", self.standard_id)
\ No newline at end of file
"""IND-006-b 邮编格式
参考:
- GB/T 23705-2009《地名 邮政编码》
6 位数字编码,前 2 位对应省级行政区域(与 GB/T 2260 一致)
- 不做「是否存在性」校验(邮编变动频繁,且新区域可能尚未公开),
仅做 6 位数字格式校验
"""
from __future__ import annotations
import re
from .base import BaseStandard, ValidationResult
class PostalCodeFormatIndicator(BaseStandard):
standard_id = "IND-006-b"
standard_name = "GB/T 23705 邮政编码 格式"
applies_to_fields = [
"postal_code", "postcode", "zip", "zip_code", "zipcode",
]
comment_keywords = ["邮编", "邮政编码"]
description = "6 位数字(GB/T 23705-2009);前 2 位对应省级区域(仅校验格式,不查存在)"
group = "通用"
severity = "warning"
_REGEX = re.compile(r"^\d{6}$")
def validate(self, value: str) -> ValidationResult:
if value is None or str(value) == "":
return ValidationResult(True, "", value, "空值跳过", self.standard_id)
v = str(value).strip()
if not self._REGEX.match(v):
return ValidationResult(
False, "", value,
f"长度 {len(v)} 或格式不符合 6 位数字",
self.standard_id,
)
return ValidationResult(True, "", value, "", self.standard_id)
\ No newline at end of file
"""IND-007-a 个人公积金账号 格式(仅位数校验)
参考:
- 公积金个人账号各省口径不一:
- 北京 / 上海 / 广州:18 位
- 深圳:10 位(早期为 11 位)
- 部分省市混合:10 位 / 12 位 / 18 位 均有
- 各省未统一发部颁格式,业内做法:**只校验位数 + 全数字**
- 不做存在性 / 校验位(无公开算法)
本 indicator 严格按用户要求:「仅实现位数 check」
- 全数字
- 位数 ∈ {10, 12, 18}
- 长度不在白名单 → 视为异常
注意:本 indicator 不查字段名(如 "fund_account" / "gjjzh"),**只查字段注释**。
原因:公积金账号字段名在不同业务系统命名差异极大(fund_acc / gjj_account /
provident_fund_no / 个人公积金账号 ...),注释反而稳定。
"""
from __future__ import annotations
import re
from .base import BaseStandard, ValidationResult
class FundAccountFormatIndicator(BaseStandard):
standard_id = "IND-007-a"
standard_name = "个人公积金账号 格式(位数)"
applies_to_fields = [] # 故意留空 —— 走字段注释匹配
comment_keywords = [
"公积金账号", "个人公积金", "公积金个人账号", "住房公积金账号",
"个人公积金账号", "公积金编号", "公积金账户", "个人公积金账户",
]
description = (
"仅位数 check:全数字,长度 ∈ {10, 12, 18};"
"各省口径不一(10/12/18 位都合法)—— 不查存在性 / 校验位(无公开算法);"
"**字段匹配走注释**,字段名差异大、不稳定"
)
group = "通用"
severity = "warning"
_REGEX = re.compile(r"^\d+$")
_ALLOWED_LENS = (10, 12, 18)
def validate(self, value: str) -> ValidationResult:
if value is None or str(value) == "":
return ValidationResult(True, "", value, "空值跳过", self.standard_id)
v = str(value).strip()
if not self._REGEX.match(v):
return ValidationResult(
False, "", value,
"非全数字字符(含字母 / 符号)",
self.standard_id,
)
if len(v) not in self._ALLOWED_LENS:
return ValidationResult(
False, "", value,
f"长度 {len(v)} ∉ {self._ALLOWED_LENS}(公积金账号通常 10 / 12 / 18 位)",
self.standard_id,
)
return ValidationResult(True, "", value, "", self.standard_id)
\ No newline at end of file
"""IND-008-a 首次参加工作年月 格式
参考:
- 劳动部门口径:YYYYMM 6 位数字编码(如 199508 表示 1995 年 8 月参加工作)
- 部分系统录入为 8 位 YYYYMMDD(带日)—— 本 indicator 接受 6 / 8 位两种
- 月份 ∈ [01, 12]
- 年份 ∈ [1940, 当前年 + 1](极端兜底:极少业务录入未来年月)
样例:
✓ 199508 / 19950808 / 202003
✗ 1995-08 / 199513 / 189908 / 203008
"""
from __future__ import annotations
import re
from .base import BaseStandard, ValidationResult
class FirstWorkMonthFormatIndicator(BaseStandard):
standard_id = "IND-008-a"
standard_name = "首次参加工作年月 格式"
applies_to_fields = [
"first_work_date", "first_work_month", "work_start_date",
"first_job_date", "ccgzrq", "scgzrq",
]
comment_keywords = [
"首次参加工作", "首次工作", "参加工作年月", "参加工作时间",
"入职时间", "入职年月", "工龄起算", "起始工作时间",
"首次参加工作日期", "首次工作日期",
]
description = (
"6 位 YYYYMM 或 8 位 YYYYMMDD;月份 ∈ [01,12];"
"年份 ∈ [1940, 当前年+1](劳动部门通用口径)"
)
group = "通用"
severity = "warning"
_REGEX_6 = re.compile(r"^(19|20)\d{2}(0[1-9]|1[0-2])$")
_REGEX_8 = re.compile(r"^(19|20)\d{2}(0[1-9]|1[0-2])(0[1-9]|[12]\d|3[01])$")
@staticmethod
def _parse_year_month(v: str) -> tuple[int, int] | None:
"""从 6 / 8 位字符串解出 (年, 月)。"""
if len(v) == 6 and v.isdigit():
try:
return int(v[:4]), int(v[4:6])
except ValueError:
return None
if len(v) == 8 and v.isdigit():
try:
return int(v[:4]), int(v[4:6])
except ValueError:
return None
return None
def validate(self, value: str) -> ValidationResult:
if value is None or str(value) == "":
return ValidationResult(True, "", value, "空值跳过", self.standard_id)
v = str(value).strip()
# 允许常见分隔形式:1995-08 / 1995.08 / 1995/08 —— 转 6 位后再校验
cleaned = re.sub(r"[\-./]", "", v)
if len(cleaned) in (6, 8) and cleaned.isdigit():
ok = bool(self._REGEX_6.match(cleaned) or self._REGEX_8.match(cleaned))
if not ok:
return ValidationResult(
False, "", value,
f"YYYYMM/YYYYMMDD 数值不合法:{cleaned}(月份必须 ∈ [01,12],日期必须 ∈ [01,31])",
self.standard_id,
)
elif not (self._REGEX_6.match(v) or self._REGEX_8.match(v)):
return ValidationResult(
False, "", value,
f"不符合 6 位 YYYYMM 或 8 位 YYYYMMDD(清洗后 = {cleaned})",
self.standard_id,
)
# 年份上限校验
ym = self._parse_year_month(cleaned)
if ym is None:
return ValidationResult(True, "", value, "", self.standard_id)
from datetime import datetime
cur_year = datetime.now().year
year, month = ym
if year < 1940:
return ValidationResult(False, "", value, f"年份 {year} < 1940(疑似异常)", self.standard_id)
if year > cur_year + 1:
return ValidationResult(False, "", value, f"年份 {year} > 当前年+1(疑似未来录入)", self.standard_id)
if month < 1 or month > 12:
return ValidationResult(False, "", value, f"月份 {month} ∉ [01,12]", self.standard_id)
return ValidationResult(True, "", value, "", self.standard_id)
\ No newline at end of file
"""IND-009-a 用工类型 值域
参考:
- 人社部《劳动合同法》口径 + 公积金行业惯例:用工类型核心枚举
01 = 正式职工(劳动合同用工)
02 = 劳务派遣(派遣员工)
03 = 灵活就业(个体工商户 / 自由职业)
04 = 实习生(见习生)
05 = 临时工
99 = 其他
- 不同省市枚举值可能略有差异,但以上 6 值覆盖 95% 业务场景
样例:
✓ 01 / 02 / 03 / 04 / 05 / 99
正式 / 劳务派遣 / 灵活就业 / 实习生 / 临时 / 其他
✗ 06 / 100 / abc / 空字符 / 0
"""
from __future__ import annotations
from .base import BaseStandard, ValidationResult
class EmploymentTypeFormatIndicator(BaseStandard):
standard_id = "IND-009-a"
standard_name = "用工类型 值域"
applies_to_fields = [
"employment_type", "employment_form", "employment_nature",
"yglx", "ygxz",
]
comment_keywords = [
"用工类型", "用工形式", "用工性质", "就业类型", "就业形式",
"劳动合同类型", "用工方式",
]
description = (
"枚举值 ∈ {01 正式 / 02 劳务派遣 / 03 灵活就业 / 04 实习生 / 05 临时工 / 99 其他};"
"枚举值不一致的录入视为异常(业务上无法分类)"
)
group = "通用"
severity = "warning"
# 同时接受数字码("01"/"02"/...)和文字("正式"/"劳务派遣"/...)
_ALLOWED = {
# 数字码
"01", "02", "03", "04", "05", "99",
# 文字形式(兜底,常见汉字录入)
"正式", "正式工", "正式职工",
"劳务派遣", "派遣", "派遣工",
"灵活就业", "灵活", "个体", "个体工商户", "自由职业",
"实习", "实习生", "见习", "见习生",
"临时", "临时工",
"其他",
}
def validate(self, value: str) -> ValidationResult:
if value is None or str(value) == "":
return ValidationResult(True, "", value, "空值跳过", self.standard_id)
v = str(value).strip()
if v in self._ALLOWED:
return ValidationResult(True, "", value, "", self.standard_id)
# 容错:去「工」/「员」/「职工」等后缀再试一次
compact = v.replace("职工", "").replace("员工", "").strip()
if compact in self._ALLOWED:
return ValidationResult(True, "", value, "", self.standard_id)
return ValidationResult(
False, "", value,
f"用工类型 {v!r} 不在白名单({sorted(self._ALLOWED)[:8]} ...)",
self.standard_id,
)
\ No newline at end of file
"""IND-010-a 本地户籍标记
参考:
- 公积金 / 社保业务普遍用布尔量:表示「是否本市户籍」/「是否本地户籍」/
「是否本市常住」
- 字段存值五花八门,需统一接受:
- 数字:1 / 0(最常见)
- 文字:是 / 否(中等常见)
- 字母:Y / N、T / F(最少见,境外业务才有)
- 不接受负数 / 小数 / 多位数字 / 空(除非表定义允许 NULL —— 空值由上游过滤)
样例:
✓ 1 / 0 / 是 / 否 / Y / N / TRUE / FALSE / true / false
✗ 2 / 1.5 / 01 / 10 / abc
"""
from __future__ import annotations
from .base import BaseStandard, ValidationResult
class LocalHouseholdFlagIndicator(BaseStandard):
standard_id = "IND-010-a"
standard_name = "本地户籍标记 值域"
applies_to_fields = [
"local_household", "is_local", "local_resident",
"bdsf", "bdshj", "sfbds",
]
comment_keywords = [
"本地户籍", "本市户籍", "本市常住", "本地常住",
"户籍本地", "本地户口", "本市户口", "户籍地",
"是否本地", "是否本市",
]
description = (
"布尔量:1/0、是/否、Y/N、T/F、TRUE/FALSE;"
"不接受 2 / 1.5 / 01 / 10 等业务上不可解释的值"
)
group = "通用"
severity = "warning"
_ALLOWED = {
"1", "0",
"是", "否",
"Y", "N", "y", "n",
"T", "F", "t", "f",
"TRUE", "FALSE", "True", "False", "true", "false",
"yes", "no", "Yes", "No", "YES", "NO",
}
def validate(self, value: str) -> ValidationResult:
if value is None or str(value) == "":
return ValidationResult(True, "", value, "空值跳过", self.standard_id)
v = str(value).strip()
if v in self._ALLOWED:
return ValidationResult(True, "", value, "", self.standard_id)
return ValidationResult(
False, "", value,
f"本地户籍标记 {v!r} 不在白名单(1/0、是/否、Y/N、T/F)",
self.standard_id,
)
\ No newline at end of file
"""IND-011-a 职工状态 值域 + 单行流转一致性
参考:
- 公积金业务口径:职工状态核心三态
01 = 正常(正常缴存中)
02 = 封存(暂停缴存但账户保留;可偿还贷款、不可办理新缴存)
03 = 销户(账户已注销;所有缴存 / 贷款业务都已结束)
04 = 转出(账户已转移到其他公积金中心)
99 = 其他
- 常见文字形式:正常 / 封存 / 销户 / 转出
流转限制(业务规则,本 indicator 仅做单行提示,跨行校验需后续 SQL/ETL):
1. 状态=销户:账户已关闭,不应再有「缴存 / 贷款 / 提取」等业务流水
—— 应核对缴存表 / 贷款表中无此账户的活动记录
2. 状态=封存:账户保留但暂停缴存,不应发起「缴存」业务;
但「偿还贷款」属于贷后行为,允许
3. 状态=正常:可缴存 / 可提取 / 可贷款 —— 无限制
本 indicator 仅做值域校验 + 单行提示(说明流转应核对的业务表);
真正的「销户后又有缴存」类违规需要跨表 SQL 关联,留待后续 indicator 实现。
样例:
✓ 01 / 02 / 03 / 04 / 99 / 正常 / 封存 / 销户 / 转出
✗ 05 / 100 / abc
"""
from __future__ import annotations
from .base import BaseStandard, ValidationResult
class EmployeeStatusFormatIndicator(BaseStandard):
standard_id = "IND-011-a"
standard_name = "职工状态 值域"
applies_to_fields = [
"employee_status", "staff_status", "worker_status",
"account_status", "zgzt", "zhzt",
]
comment_keywords = [
"职工状态", "员工状态", "账户状态", "公积金状态",
"人员状态", "缴存状态", "参保状态",
]
description = (
"枚举值 ∈ {01 正常 / 02 封存 / 03 销户 / 04 转出 / 99 其他};\n"
"流转限制(业务规则,跨表 SQL 校验):\n"
" - 状态=销户:账户已关闭,不应再有缴存 / 贷款 / 提取流水\n"
" - 状态=封存:账户保留但暂停缴存,不应发起缴存;可偿还贷款\n"
" - 状态=正常:可缴存 / 提取 / 贷款\n"
"(流转限制不在本 indicator 内做 —— 留给跨表 SQL 后续实现)"
)
group = "通用"
severity = "warning"
_ALLOWED = {
# 数字码
"01", "02", "03", "04", "99",
# 文字形式(兜底常见汉字录入)
"正常", "封存", "销户", "封存停缴", "已封存", "已销户", "转出", "已转出",
"封存暂停", "正常缴存",
}
def validate(self, value: str) -> ValidationResult:
if value is None or str(value) == "":
return ValidationResult(True, "", value, "空值跳过", self.standard_id)
v = str(value).strip()
if v in self._ALLOWED:
return ValidationResult(True, "", value, "", self.standard_id)
return ValidationResult(
False, "", value,
f"职工状态 {v!r} 不在白名单(01 正常 / 02 封存 / 03 销户 / 04 转出 / 99 其他)",
self.standard_id,
)
\ No newline at end of file
......@@ -74,6 +74,7 @@ def list_all_standards() -> list[dict]:
"id": inst.standard_id,
"name": inst.standard_name,
"applies_to_fields": inst.applies_to_fields,
"comment_keywords": getattr(inst, "comment_keywords", []) or [],
"description": inst.description,
"group": getattr(inst, "group", "通用"),
"severity": getattr(inst, "severity", "warning"),
......
......@@ -16,7 +16,7 @@
{
"key": "standards",
"title": "国标字段规范",
"description": "IND-001 ~ IND-004 系列:按 GB / GA 标准做字段值级别合规校验(格式 / 校验位 / 出生日期 / 15 位兼容 / 号段 / 编码存在性)",
"description": "IND-001 ~ IND-005 系列:按 GB / GA 标准做字段值级别合规校验(格式 / 校验位 / 出生日期 / 15 位兼容 / 号段 / 编码存在性 / 固定电话)",
"default_expand": true,
"children": [
{ "step_id": "std_ind_001_a" },
......@@ -28,7 +28,8 @@
{ "step_id": "std_ind_003_a" },
{ "step_id": "std_ind_003_b" },
{ "step_id": "std_ind_004_a" },
{ "step_id": "std_ind_004_b" }
{ "step_id": "std_ind_004_b" },
{ "step_id": "std_ind_005_a" }
]
},
{
......@@ -68,6 +69,21 @@
{ "step_id": "std_ind_602" },
{ "step_id": "std_ind_603" }
]
},
{
"key": "business",
"title": "业务字段规范",
"description": "IND-006 ~ IND-011 系列:业务特定字段(通讯地址 / 邮编 / 个人公积金账号 / 首次参加工作年月 / 用工类型 / 本地户籍 / 职工状态);字段匹配走注释",
"default_expand": true,
"children": [
{ "step_id": "std_ind_006_a" },
{ "step_id": "std_ind_006_b" },
{ "step_id": "std_ind_007_a" },
{ "step_id": "std_ind_008_a" },
{ "step_id": "std_ind_009_a" },
{ "step_id": "std_ind_010_a" },
{ "step_id": "std_ind_011_a" }
]
}
]
}
......@@ -260,6 +260,7 @@ def _derive_format(meta: dict) -> str:
def _derive_target(meta: dict) -> str:
"""从 indicator metadata 派生「查询对象」短句。"""
fields = meta.get("applies_to_fields") or []
cmts = [k for k in (meta.get("comment_keywords") or []) if k]
sid = meta.get("id", "")
# 跨字段 / 跨表 indicator:title 已暗示了
if sid in ("IND-301",):
......@@ -280,6 +281,28 @@ def _derive_target(meta: dict) -> str:
return "字段名含 marriage / hyzk / marital 等婚姻状况字段"
if sid in ("IND-101",):
return "字段名含 id_type / zjlx / cert_type 等证件类型字段"
# 公积金 / 业务领域 indicator:comment-only 命中
if sid in ("IND-005-a",):
return "字段名或注释含「固定电话 / 办公电话 / 公司电话 / 传真 / 座机」"
if sid in ("IND-006-a",):
return "字段名或注释含「通讯地址 / 居住地址 / 户籍地址 / 公司地址」"
if sid in ("IND-006-b",):
return "字段名或注释含「邮编 / 邮政编码」"
if sid in ("IND-007-a",):
return "**字段注释**含「公积金账号 / 个人公积金 / 公积金账户」"
if sid in ("IND-008-a",):
return "字段名或注释含「首次参加工作 / 参加工作年月 / 入职时间」"
if sid in ("IND-009-a",):
return "字段名或注释含「用工类型 / 用工形式 / 就业类型」"
if sid in ("IND-010-a",):
return "字段名或注释含「本地户籍 / 本市户籍 / 是否本地」"
if sid in ("IND-011-a",):
return "字段名或注释含「职工状态 / 账户状态 / 缴存状态」"
# 兜底:优先显示 comment 命中关键字,否则按字段名
if cmts and not fields:
return "**字段注释**含 " + " / ".join(cmts)
if cmts and fields:
return "字段名含 " + " / ".join(fields[:6]) + "(或注释含 " + " / ".join(cmts[:4]) + ")"
if fields:
return "字段名含 " + " / ".join(fields)
return "(跨字段 / 跨表,见 indicator 名称)"
......@@ -306,10 +329,41 @@ def _derive_check(meta: dict) -> str:
"IND-001-d": "15 位老证:长度 = 15,6 位地址 + 6 位出生日期(无世纪位)+ 3 位顺序码;按 19xx 补齐后校验。",
"IND-002-a": "正则 ^ [0-9A-HJ-NPQRTUWXY]{2}\\d{6}[0-9A-HJ-NPQRTUWXY]{10}$(18 位,31 字符集不含 I/O/Z/S/V)。",
"IND-002-b": "GB 32100-2015 MOD 31-3:前 17 位 × 权重 [3,7,9,10,5,8,4,2,...],模 31,查映射表(含 0-9 + 25 个字母)得校验位。",
"IND-003-a": "11 位数字 + 第 1 位 1[3-9] / 第 2 位 3-9;正则 ^1[3-9]\\d{9}$。",
"IND-003-a": (
"1) 自动清洗:去 +86 / 86 前缀、去空格 / 横线\n"
"2) 清洗后长度 = 11\n"
"3) 头部 1[3-9]:第 1 位 = 1、第 2 位 ∈ [3-9]\n"
"4) **不可全 0**(00000000000 等号段残缺值视为无效)"
),
"IND-003-b": "查工信部公开号段表(170/171/180-189 等),前缀必须在号段内;号段字典每季度更新。",
"IND-004-a": "6 位数字:[省 2 位][地市 2 位][县区 2 位];首位非 0;不强制存在性(IND-004-b 才校验)。",
"IND-004-b": "查 GB/T 2260-2007 表(约 4400 条),区码必须存在;港澳台前缀 71/81/82;用户表中 83 也接受(GB/T 2260 不含但 GB/T 22239 引用)。",
"IND-005-a": "正则 ^0\\d{2,3}[-\\s]?\\d{7,8}$;区号 3~4 位(必 0 开头)+ 可选「- / 空格」+ 号码 7~8 位;不查号段(工信部未公开固话号段表)。",
"IND-006-a": (
"1) trim 后长度 ∈ [4, 200]\n"
"2) 不可全空白 / 全符号 / 全数字\n"
"3) 不为「无 / 暂无 / 未知 / TBD」占位符\n"
"4) 不做行政区划 / 路名 / 门牌号语义(无国标)"
),
"IND-006-b": "正则 ^\\d{6}$;前 2 位对应省级行政区域(与 GB/T 2260 一致);仅校验格式,不查邮编存在性(GB/T 23705-2009 表变动频繁)。",
"IND-007-a": (
"仅位数 check:\n"
" 1) 全数字\n"
" 2) 长度 ∈ {10, 12, 18}(各省口径不一)\n"
"不查存在性 / 校验位(无公开算法)\n"
"字段匹配走注释,applies_to_fields 故意留空"
),
"IND-008-a": "6 位 YYYYMM 或 8 位 YYYYMMDD;月份 ∈ [01,12];年份 ∈ [1940, 当前年+1];分隔符(- . /)清洗后再校验。",
"IND-009-a": "枚举值 ∈ {01 正式 / 02 劳务派遣 / 03 灵活就业 / 04 实习生 / 05 临时工 / 99 其他};接受数字码 + 文字形式(含「工 / 职工 / 员工」后缀容错)。",
"IND-010-a": "布尔量:1/0、是/否、Y/N、T/F、TRUE/FALSE;不接受 2 / 1.5 / 01 / 10 等业务上不可解释的值。",
"IND-011-a": (
"1) 状态值域 ∈ {01 正常 / 02 封存 / 03 销户 / 04 转出 / 99 其他}(接受数字码 + 文字形式)\n"
"2) 流转限制(业务规则,跨表 SQL 校验):\n"
" · 销户:账户已关闭,不应再有缴存 / 贷款 / 提取流水\n"
" · 封存:账户保留但暂停缴存;不应发起缴存;可偿还贷款\n"
" · 正常:可缴存 / 提取 / 贷款\n"
"流转校验本 indicator 不做 —— 留给跨表 SQL 后续实现"
),
"IND-101": "5 值枚举:01=身份证 / 03=护照 / 04=永居证 / 05=港澳台居住证 / 99=其他(用户给定的简化集,未走 GA/T 2000.156-2016)。",
"IND-201": "复用 IND-001-a/b/c 的 18 位 GB 11643 + 校验位 + 出生日期校验(与 IND-001 是有意重叠;按业务语义选用)。",
"IND-202": "9 位:G+8 / E+8 / E+1字母+7数字 / PE/SE/DE+7数字;字母排除 I/O;无公开校验位(只做格式校验)。",
......@@ -357,8 +411,16 @@ def _register_indicator_steps() -> None:
same_group_count += 1
order = base_order + same_group_count
title = f"{std_id} · {meta['name']}"
fields = meta.get("applies_to_fields") or []
cmts = [k for k in (meta.get("comment_keywords") or []) if k]
match_desc = []
if fields:
match_desc.append("字段名: " + ", ".join(fields[:8]) + (" ..." if len(fields) > 8 else ""))
if cmts:
match_desc.append("注释关键字: " + ", ".join(cmts))
match_summary = " / ".join(match_desc) or "(跨字段/跨表)"
description = (
f"分组: {group} | 适用字段: {', '.join(meta.get('applies_to_fields', [])) or '(跨字段/跨表)'} | "
f"分组: {group} | 匹配: {match_summary} | "
f"{meta.get('description', '')}"
)
detail = StepDetail(
......
......@@ -462,9 +462,58 @@ def _empty_indicator_data() -> dict:
}
def _collect_hit_columns(
std_instance: BaseStandard,
all_columns: list[dict],
) -> tuple[list[dict], str]:
"""按 applies_to_fields(字段名) + comment_keywords(字段注释 substring,
不区分大小写)两路并集命中字段;返回 (命中字段列表, 命中原因描述)。
"""
applicable = list(getattr(std_instance, "applies_to_fields", []) or [])
comment_kws = [k.strip().lower() for k in (getattr(std_instance, "comment_keywords", []) or []) if k]
applicable_set = set(applicable)
hits: list[dict] = []
seen: set[tuple[str, str]] = set()
hit_kinds: set[str] = set()
# 第 1 路:按字段名
if applicable_set:
for c in all_columns:
col_name = c.get("column_name", "")
if col_name in applicable_set:
key = (c.get("table_name", ""), col_name)
if key not in seen:
seen.add(key)
hits.append({**c, "_hit_kind": "name", "_hit_key": col_name})
hit_kinds.add("name")
# 第 2 路:按字段注释(substring 不区分大小写)
if comment_kws:
for c in all_columns:
cc = (c.get("column_comment") or "").lower()
if not cc:
continue
for kw in comment_kws:
if kw and kw in cc:
key = (c.get("table_name", ""), c.get("column_name", ""))
if key not in seen:
seen.add(key)
hits.append({**c, "_hit_kind": "comment", "_hit_key": kw})
hit_kinds.add("comment")
break
reason_bits = []
if "name" in hit_kinds:
reason_bits.append(f"字段名匹配 {applicable}")
if "comment" in hit_kinds:
reason_bits.append(f"注释匹配 {comment_kws}")
return hits, " / ".join(reason_bits) or "(无命中条件)"
def _run_round1_one(
cfg: DBConfig,
by_name: dict[str, list[dict]],
all_columns: list[dict],
std_instance: BaseStandard,
log: Callable | None,
) -> dict:
......@@ -482,22 +531,36 @@ def _run_round1_one(
f" · {std_id} ({getattr(std_instance, 'group', '通用')}) 启动",
step=indicator_step_id(std_id))
# 找该 indicator 命中的所有字段(applies_to_fields 中任意一个)
applicable = list(getattr(std_instance, "applies_to_fields", []) or [])
# 找该 indicator 命中的所有字段(applies_to_fields ∪ comment_keywords)
hit_cols, hit_reason = _collect_hit_columns(std_instance, all_columns)
# 单 (table, field) 命中后仍按 MAX_TABLES_PER_FIELD 截断同一字段的表数;
# 不同表名 → 不同字段对 → 都保留
fields_to_check: list[dict] = []
for f in applicable:
if f in by_name:
for c in by_name[f][:MAX_TABLES_PER_FIELD]:
# 标记命中的字段名(一个 indicator 可能命中多列)
fields_to_check.append({**c, "_hit_field": f})
per_field_tables: dict[str, int] = {}
for c in hit_cols:
col_name = c.get("column_name", "")
per_field_tables[col_name] = per_field_tables.get(col_name, 0)
if per_field_tables[col_name] >= MAX_TABLES_PER_FIELD:
continue
per_field_tables[col_name] += 1
fields_to_check.append(c)
if not fields_to_check:
applicable = list(getattr(std_instance, "applies_to_fields", []) or [])
cmt_kws = [k for k in (getattr(std_instance, "comment_keywords", []) or []) if k]
if log:
log("DEBUG",
f" · {std_id} 未匹配任何字段(适用 {applicable}),跳过",
f" · {std_id} 未匹配任何字段(字段名={applicable} / "
f"注释关键字={cmt_kws}),跳过",
step=indicator_step_id(std_id))
return bucket
if log:
log("DEBUG",
f" · {std_id} 命中 {len(fields_to_check)} 个字段({hit_reason})",
step=indicator_step_id(std_id))
rec = {
"indicator": std_id,
"indicator_name": std_instance.standard_name,
......@@ -624,7 +687,7 @@ def _run_round1_one(
def _run_round1(
cfg: DBConfig,
by_name: dict[str, list[dict]],
all_columns: list[dict],
standards_by_field: dict[str, BaseStandard],
log: Callable | None,
) -> dict:
......@@ -649,11 +712,11 @@ def _run_round1(
if log:
log("INFO",
f"Round 1 启动: 共 {len(seen_instances)} 个 indicator, "
f"按 applies_to_fields 命中 {len(standards_by_field)} 个字段",
f"扫描 {len(all_columns)} 个字段",
step="standards")
for std_id, std_instance in seen_instances.items():
single = _run_round1_one(cfg, by_name, std_instance, log)
single = _run_round1_one(cfg, all_columns, std_instance, log)
group = getattr(std_instance, "group", "通用")
bucket = per_group.setdefault(group, _empty_group_data())
# 合并:violations_by_indicator + violations
......@@ -1004,12 +1067,12 @@ def run_step7(cfg: DBConfig, dict_data: dict, log: Callable | None = None,
if log:
log("INFO",
f"已加载 {len(all_standards_info)} 个标准插件, "
f"匹配字段 {len(standards_by_field)} 个, "
f"扫描字段 {len(columns)} 个(按字段名 / 注释双路匹配), "
f"最大每字段抽样 {SAMPLE_LIMIT} 行 / 每字段最多 {MAX_TABLES_PER_FIELD} 张表",
step="standards")
# ── Round 1 单值校验 ──
per_group = _run_round1(cfg, by_name, standards_by_field, log)
# ── Round 1 单值校验(按 applies_to_fields ∪ comment_keywords 命中)──
per_group = _run_round1(cfg, columns, standards_by_field, log)
# ── Round 2 跨字段(IND-301)──
cross_field = _run_round2_cross_field(cfg, by_name, log)
......@@ -1077,10 +1140,6 @@ def run_step7_for_indicator(
std_instance = cls()
by_name: dict[str, list[dict]] = {}
for r in columns:
by_name.setdefault(r["column_name"], []).append(r)
if not columns:
return _empty_indicator_step_result(standard_id)
......@@ -1089,19 +1148,27 @@ def run_step7_for_indicator(
safe_id = step_id_str[len("std_"):] # "ind_001_a"
section_key = f"standards_{safe_id}" # "standards_ind_001_a"
# IND-301/302 是按 (table, col) 列名配对的,仍走老 by_name 入口
by_name: dict[str, list[dict]] = {}
for r in columns:
by_name.setdefault(r["column_name"], []).append(r)
if standard_id == "IND-301":
data = _run_round2_cross_field_one(cfg, by_name, log)
elif standard_id == "IND-302":
data = _run_round2_uniqueness_one(cfg, by_name, log)
else:
# 单值类(含 22 个有 applies_to_fields 的 indicator)
if not std_instance.applies_to_fields:
# 单值类:字段名 OR 注释命中都可(applies_to_fields / comment_keywords 二选一非空)
has_name = bool(getattr(std_instance, "applies_to_fields", None))
has_cmt = bool(getattr(std_instance, "comment_keywords", None))
if not (has_name or has_cmt):
if log:
log("WARN",
f"{standard_id} 没有 applies_to_fields 也没匹配 IND-301/302 模板",
f"{standard_id} 没有 applies_to_fields / comment_keywords 也没匹配 "
f"IND-301/302 模板,跳过",
step=indicator_step_id(standard_id))
return _empty_indicator_step_result(standard_id)
data = _run_round1_one(cfg, by_name, std_instance, log)
data = _run_round1_one(cfg, columns, std_instance, log)
return {"section_key": section_key, "data": _wrap_single_indicator(std_instance, data)}
......
......@@ -443,50 +443,58 @@ body {
.log-error .log-level { color: #f48771; }
.log-error { background: rgba(244, 135, 113, 0.1); }
/* ── KPI 卡片 ── */
/* ── KPI ── */
.kpi-grid {
display: grid;
grid-template-columns: repeat(auto-fit, minmax(140px, 1fr));
gap: 16px;
grid-template-columns: repeat(auto-fit, minmax(120px, 1fr));
gap: 8px 24px; /* 上下紧凑 / 左右放宽,密排不拥挤 */
margin-bottom: 20px;
align-items: baseline;
}
.kpi-card {
background: white;
border-radius: 8px;
padding: 16px;
text-align: center;
box-shadow: 0 2px 8px rgba(0, 0, 0, 0.06);
border-left: 4px solid #1890ff;
transition: all 0.2s;
/* 去掉卡片感 —— 不要背景 / 阴影 / 边框 / 内边距 */
text-align: left;
transition: color 0.15s;
}
.kpi-card:hover {
transform: translateY(-2px);
box-shadow: 0 4px 12px rgba(0, 0, 0, 0.1);
.kpi-card.kpi-tables .kpi-value,
.kpi-card.kpi-fields .kpi-value,
.kpi-card.kpi-empty .kpi-value,
.kpi-card.kpi-empty-mid .kpi-value,
.kpi-card.kpi-missing .kpi-value,
.kpi-card.kpi-length .kpi-value,
.kpi-card.kpi-std .kpi-value,
.kpi-card.kpi-merge .kpi-value,
.kpi-card.kpi-redundancy .kpi-value {
/* 保留颜色 hint,但用文字色而非左边框 */
color: #1890ff;
}
.kpi-card.kpi-tables { border-left-color: #1890ff; }
.kpi-card.kpi-fields { border-left-color: #13c2c2; }
.kpi-card.kpi-empty { border-left-color: #faad14; }
.kpi-card.kpi-empty-mid { border-left-color: #e6a23c; }
.kpi-card.kpi-missing { border-left-color: #fa8c16; }
.kpi-card.kpi-length { border-left-color: #f5222d; }
.kpi-card.kpi-std { border-left-color: #eb2f96; }
.kpi-card.kpi-merge { border-left-color: #722ed1; }
.kpi-card.kpi-redundancy { border-left-color: #52c41a; }
.kpi-card.kpi-tables .kpi-value { color: #1890ff; }
.kpi-card.kpi-fields .kpi-value { color: #13c2c2; }
.kpi-card.kpi-empty .kpi-value { color: #faad14; }
.kpi-card.kpi-empty-mid .kpi-value { color: #e6a23c; }
.kpi-card.kpi-missing .kpi-value { color: #fa8c16; }
.kpi-card.kpi-length .kpi-value { color: #f5222d; }
.kpi-card.kpi-std .kpi-value { color: #eb2f96; }
.kpi-card.kpi-merge .kpi-value { color: #722ed1; }
.kpi-card.kpi-redundancy .kpi-value { color: #52c41a; }
.kpi-value {
font-size: 32px;
font-size: 28px;
font-weight: 700;
color: #303133;
color: #1890ff;
line-height: 1.2;
font-variant-numeric: tabular-nums;
}
.kpi-label {
color: #909399;
font-size: 13px;
margin-top: 4px;
font-size: 12px;
margin-top: 2px;
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
}
.muted {
......
Markdown is supported
0%
or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment