Commit a72a2ec1 authored by Data Governance Dev's avatar Data Governance Dev

feat(standards): IND-002/003/004 合并 + IND-005 改名(去 -a)

需求:把 IND-002(USCC 3 子)/ IND-003(手机号 2 子)/ IND-004(行政区划 2 子)
也和 IND-001 一样合并成单 step;IND-005(固定电话,1 子)UI 不再显示 -a 后缀。

改动:
- step7_standards.py: 把 IND-001 专属的合并逻辑通用化为配置驱动
  - _IND_COMBINED: {parent_id: {sub_ids, title, format}} 配置表(覆盖 ind_001/-002/-003/-004)
  - _wrap_combined_ind_001 → _wrap_combined_indicator(parent_id, ...)
  - run_step7_for_ind_001 → run_step7_combined(parent_id, ...)
  - 旧名保留为薄包装(backward compat)
- orchestrator.py:
  - _MERGED_SUBS 扩展到 12 个子指标(4+3+2+2+1)
  - _run_standards_combined(parent_id) 工厂 + _COMBINED_STEP_META 注册表
  - 循环注册 std_ind_001/-002/-003/-004 四个 step
  - _run_standards_ind_005 端到端改 4 处:section_key / step_id / tab.key / tab.title
- analysis_tree.json: 标准组 7 行 → 5 行(去掉 -a/-b/-c 后缀)
- routes.py: _COMBINED_GROUPS 4 组聚合 + IND-005 单独处理(YAML key 仍叫 IND-005-a)
- WORKLOG 追加整段记录 + 踩坑说明(IND-005 改名必须 4 处一致)
parent ba75ed1b
...@@ -2,6 +2,86 @@ ...@@ -2,6 +2,86 @@
> 任务做完一次记一次。最近的在最上面。 > 任务做完一次记一次。最近的在最上面。
## 2026-08-13 · IND-002 / -003 / -004 合并 + IND-005 改名(去 -a)
### 需求
> 用户反馈:「同样的把 IND-002/IND-003/IND-004 也合并,IND-005 不显示 -a」
2026-08-12 只合并了 IND-001(4 子检查),其余 4 个 IND 各自的 a/b/c 子检查仍各跑各的,结果散落多个 tab。IND-005 只有 1 个子检查(IND-005-a 固定电话),但 UI 上显示了 -a 后缀视觉不一致。
### 设计决策(用户已确认)
1. **IND-002(USCC 统一社会信用代码)**:3 子检查 → 合并到 `std_ind_002`
2. **IND-003(手机号)**:2 子检查 → 合并到 `std_ind_003`
3. **IND-004(行政区划代码)**:2 子检查 → 合并到 `std_ind_004`
4. **IND-005(固定电话)**:仅 1 子检查,**改 step_id + 标题去掉 -a 后缀**(`std_ind_005_a` → `std_ind_005`)
### 改动 / 新增的 3 个文件
**[web/core/step_impl/step7_standards.py](web/core/step_impl/step7_standards.py)** — 把 IND-001 专属的合并逻辑**通用化**为配置驱动
```python
_IND_COMBINED: dict[str, dict] = {
"ind_001": {"sub_ids": ("IND-001-a", "-b", "-c", "-d"),
"title": "IND-001 · 身份证号校验",
"format": "GB 11643-1999(合并 a/b/c/d)/ GB 11643-1989(d 子项兼容)"},
"ind_002": {"sub_ids": ("IND-002-a", "-b", "-c"),
"title": "IND-002 · 统一社会信用代码校验",
"format": "GB 32100-2015(含 /-c 老代码兼容转换)"},
"ind_003": {"sub_ids": ("IND-003-a", "-b"),
"title": "IND-003 · 手机号校验",
"format": "工信部(《电信网编号计划》)"},
"ind_004": {"sub_ids": ("IND-004-a", "-b"),
"title": "IND-004 · 行政区划代码校验",
"format": "GB/T 2260"},
}
# 通用化 _wrap_combined_ind_001 → _wrap_combined_indicator(parent_id, ...)
# 通用化 run_step7_for_ind_001 → run_step7_combined(parent_id, ...)
# 旧名保留为薄包装,留作 backward compat
```
**[web/core/orchestrator.py](web/core/orchestrator.py)** — 3 大块:
1. `_MERGED_SUBS` 集合扩展到 12 个子指标(4+3+2+2+1)
2. `_run_standards_combined(parent_id)` 工厂 + `_COMBINED_STEP_META` 注册表,循环注册 4 个 step
3. `_run_standards_ind_005` 单独实现:跑 IND-005-a 单 indicator,**端到端重命名** section_key / step_id / tab.key / tab.title
**[web/configs/analysis_tree.json](web/configs/analysis_tree.json)** — 国标字段规范组 5 行:
```json
{ "step_id": "std_ind_001" },
{ "step_id": "std_ind_002" },
{ "step_id": "std_ind_003" },
{ "step_id": "std_ind_004" },
{ "step_id": "std_ind_005" }
```
**[web/api/routes.py](web/api/routes.py)** — `/api/match-config` 聚合块泛化:
```python
_COMBINED_GROUPS = (
("std_ind_001", ("IND-001-a", "IND-001-b", "IND-001-c", "IND-001-d")),
("std_ind_002", ("IND-002-a", "IND-002-b", "IND-002-c")),
("std_ind_003", ("IND-003-a", "IND-003-b")),
("std_ind_004", ("IND-004-a", "IND-004-b")),
)
# + IND-005 单独处理:YAML key 仍叫 IND-005-a,UI step_id 改为 std_ind_005
```
### 踩坑
- **IND-005 改名必须 4 处一致**:`section_key` / `_protocol.step_id` / `tab.key` / `tab.title`。
- 最初只改了 section_key 和 step_id,结果 tab.key 还是 "standards_ind_005_a"、tab.title 还是 "IND-005-a · …"
- 后来加 2 行 for tab 改 key + title
- **tab.title 不能直接 `parent_id.upper().replace('_', '-')`**:合并模板里有 `s.standard_name` 显示用,生成时会拼成 `IND-001 · 身份证号校验` 之类,所以必须是配置表的 `title` 字段
### 烟测
```
ind_001: section_key=standards_ind_001 step_id=std_ind_001 tab.key=standards_ind_001
ind_002: section_key=standards_ind_002 step_id=std_ind_002 tab.key=standards_ind_002
ind_003: section_key=standards_ind_003 step_id=std_ind_003 tab.key=standards_ind_003
ind_004: section_key=standards_ind_004 step_id=std_ind_004 tab.key=standards_ind_004
ind_005: section_key=standards_ind_005 step_id=std_ind_005 tab.key=standards_ind_005
```
`/api/analysis-tree` 标准组 5 个 checkbox、`/api/match-config` 5 个输入框聚合正确,零孤儿。
---
## 2026-08-12 · IND-001 合并后「结果又不显示」修复(double-wrap bug) ## 2026-08-12 · IND-001 合并后「结果又不显示」修复(double-wrap bug)
### 问题 ### 问题
......
...@@ -264,27 +264,43 @@ async def get_match_config(): ...@@ -264,27 +264,43 @@ async def get_match_config():
"skip": False, "skip": False,
} }
# ── 合并 step:std_ind_001(IND-001 a/b/c/d 聚合)─────────── # ── 合并 step:std_ind_001 / -002 / -003 / -004 聚合子检查 ──
# 4 个子 indicator 的 YAML 配置聚合到 1 个 step 输入框。 # N 个子 indicator 的 YAML 配置聚合到 1 个 step 输入框。
# 用户编辑的 override 在后端 run_step7_for_ind_001 里会下推到各子检查。 # 用户编辑的 override 在后端 run_step7_combined 里会下推到各子检查。
ind001_sub_ids = ("IND-001-a", "IND-001-b", "IND-001-c", "IND-001-d") _COMBINED_GROUPS = (
ind001_names: list[str] = [] ("std_ind_001", ("IND-001-a", "IND-001-b", "IND-001-c", "IND-001-d")),
ind001_comments: list[str] = [] ("std_ind_002", ("IND-002-a", "IND-002-b", "IND-002-c")),
ind001_seen_n: set = set() ("std_ind_003", ("IND-003-a", "IND-003-b")),
ind001_seen_c: set = set() ("std_ind_004", ("IND-004-a", "IND-004-b")),
for sub_id in ind001_sub_ids: )
for combined_step_id, sub_ids in _COMBINED_GROUPS:
agg_names: list[str] = []
agg_comments: list[str] = []
seen_n: set = set()
seen_c: set = set()
for sub_id in sub_ids:
rec = step7.get(sub_id) or {} rec = step7.get(sub_id) or {}
for n in (rec.get("applies_to_fields") or []): for n in (rec.get("applies_to_fields") or []):
if n not in ind001_seen_n: if n not in seen_n:
ind001_names.append(n) agg_names.append(n)
ind001_seen_n.add(n) seen_n.add(n)
for c in (rec.get("comment_keywords") or []): for c in (rec.get("comment_keywords") or []):
if c not in ind001_seen_c: if c not in seen_c:
ind001_comments.append(c) agg_comments.append(c)
ind001_seen_c.add(c) seen_c.add(c)
by_step["std_ind_001"] = { by_step[combined_step_id] = {
"names": ", ".join(ind001_names), "names": ", ".join(agg_names),
"comments": ", ".join(ind001_comments), "comments": ", ".join(agg_comments),
"skip": False,
}
# ── IND-005 不再带 -a 后缀:YAML key 仍叫 IND-005-a,UI step_id 改为 std_ind_005 ──
ind005_rec = step7.get("IND-005-a") or {}
ind005_atf = ind005_rec.get("applies_to_fields") or []
ind005_cmk = ind005_rec.get("comment_keywords") or []
by_step["std_ind_005"] = {
"names": ", ".join(ind005_atf),
"comments": ", ".join(ind005_cmk),
"skip": False, "skip": False,
} }
......
...@@ -15,18 +15,14 @@ ...@@ -15,18 +15,14 @@
{ {
"key": "standards", "key": "standards",
"title": "国标字段规范", "title": "国标字段规范",
"description": "IND-001 ~ IND-005 系列:按 GB / GA 标准做字段值级别合规校验(格式 / 校验位 / 出生日期 / 15 位兼容 / 号段 / 编码存在性 / 固定电话 / 老代码兼容)", "description": "IND-001 ~ IND-005 系列:按 GB / GA 标准做字段值级别合规校验。IND-001(身份证)/ IND-002(USCC)/ IND-003(手机号)/ IND-004(行政区划)多子检查已合并为单 step;IND-005(固定电话)单 sub。包含格式 / 校验位 / 出生日期 / 15 位兼容 / 号段 / 编码存在性 / 老代码兼容",
"default_expand": true, "default_expand": true,
"children": [ "children": [
{ "step_id": "std_ind_001" }, { "step_id": "std_ind_001" },
{ "step_id": "std_ind_002_a" }, { "step_id": "std_ind_002" },
{ "step_id": "std_ind_002_b" }, { "step_id": "std_ind_003" },
{ "step_id": "std_ind_002_c" }, { "step_id": "std_ind_004" },
{ "step_id": "std_ind_003_a" }, { "step_id": "std_ind_005" }
{ "step_id": "std_ind_003_b" },
{ "step_id": "std_ind_004_a" },
{ "step_id": "std_ind_004_b" },
{ "step_id": "std_ind_005_a" }
] ]
}, },
{ {
......
...@@ -540,11 +540,18 @@ def _register_indicator_steps() -> None: ...@@ -540,11 +540,18 @@ def _register_indicator_steps() -> None:
过滤: 过滤:
- 跳过 std_id 以 "STD-" 开头的已拆分旧插件(applies_to_fields=[]) - 跳过 std_id 以 "STD-" 开头的已拆分旧插件(applies_to_fields=[])
- 跳过 IND-001-a / -b / -c / -d(已被合并到 `std_ind_001` 复合 step,2026-08-12 起) - 跳过 IND-001/002/003/004 各 a/b/c/d 子指标(已合并到对应复合 step)
- 跳过 IND-005-a(已重命名为 std_ind_005,不再保留 -a 后缀)
- 保留 IND-301 / IND-302 等无 applies_to_fields 的跨字段 indicator - 保留 IND-301 / IND-302 等无 applies_to_fields 的跨字段 indicator
""" """
# IND-001 的 4 个子指标合并为一个;其余 indicator 仍按 1 个 step / 1 个 tab 注册 # 已合并 / 改名的子指标集合
_IND_001_SUBS = ("IND-001-a", "IND-001-b", "IND-001-c", "IND-001-d") _MERGED_SUBS = (
"IND-001-a", "IND-001-b", "IND-001-c", "IND-001-d",
"IND-002-a", "IND-002-b", "IND-002-c",
"IND-003-a", "IND-003-b",
"IND-004-a", "IND-004-b",
"IND-005-a", # 2026-08-13:去掉 -a 后缀,改名 std_ind_005
)
prev_group: str | None = None prev_group: str | None = None
same_group_count = 0 same_group_count = 0
...@@ -555,10 +562,19 @@ def _register_indicator_steps() -> None: ...@@ -555,10 +562,19 @@ def _register_indicator_steps() -> None:
if std_id.startswith("STD-"): if std_id.startswith("STD-"):
logger.info(f"跳过已 deprecated 旧插件: {std_id}(已被 IND-* 拆分)") logger.info(f"跳过已 deprecated 旧插件: {std_id}(已被 IND-* 拆分)")
continue continue
# 跳过 IND-001 子指标(已合并到 std_ind_001 复合 step) # 跳过已合并 / 改名的子指标
if std_id in _IND_001_SUBS: if std_id in _MERGED_SUBS:
mapped_step = {
"IND-001-a": "std_ind_001", "IND-001-b": "std_ind_001",
"IND-001-c": "std_ind_001", "IND-001-d": "std_ind_001",
"IND-002-a": "std_ind_002", "IND-002-b": "std_ind_002",
"IND-002-c": "std_ind_002",
"IND-003-a": "std_ind_003", "IND-003-b": "std_ind_003",
"IND-004-a": "std_ind_004", "IND-004-b": "std_ind_004",
"IND-005-a": "std_ind_005",
}.get(std_id, "?")
logger.info( logger.info(
f"跳过子 indicator: {std_id}(已合并到 std_ind_001 复合 step)" f"跳过子 indicator: {std_id}(已合并/改名到 {mapped_step} 复合 step)"
) )
continue continue
group = meta.get("group", "通用") group = meta.get("group", "通用")
...@@ -600,44 +616,168 @@ def _register_indicator_steps() -> None: ...@@ -600,44 +616,168 @@ def _register_indicator_steps() -> None:
_register_indicator_steps() _register_indicator_steps()
# ── IND-001 复合 step(合并 a/b/c/d 4 个子检查) ────────────── # ── IND-001 / -002 / -003 / -004 复合 step(合并子检查) ──────────────
# 2026-08-12 起:UI 上只展示 1 个 checkbox「IND-001 · 身份证号」, # 2026-08-12 起:UI 上只展示 1 个 checkbox「IND-XXX · …」,
# 后端内部跑全部 4 个子检查,结果按 (table, column, value) 合并展示, # 后端内部跑全部子检查,结果按 (table, column, value) 合并展示,
# 同一个值若同时违反多个子检查 → 合并成 1 行,error 字段拼所有子错误。 # 同一个值若同时违反多个子检查 → 合并成 1 行,error 字段拼所有子错误。
# scope 限定 IND-001,其余 a/b 后缀 indicator 暂不动。 # 2026-08-13 扩展到 IND-002(USCC)/ IND-003(手机号)/ IND-004(行政区划);
def _run_standards_ind_001(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides): # noqa: ARG001 # 共用 _IND_COMBINED 配置 + run_step7_combined 入口。
from .step_impl.step7_standards import run_step7_for_ind_001 def _run_standards_combined(parent_id: str):
# run_step7_for_ind_001 已返回 {section_key, data},直接透传(不要再 wrap) """工厂:生成 step 闭包,调用 run_step7_combined(parent_id, ...)。"""
return run_step7_for_ind_001( def _runner(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides): # noqa: ARG001
cfg, dict_data=dict_data, log=log, from .step_impl.step7_standards import run_step7_combined
# run_step7_combined 已返回 {section_key, data},直接透传(不要再 wrap)
return run_step7_combined(
parent_id, cfg, dict_data=dict_data, log=log,
table_filter=table_filter, match_overrides=match_overrides, table_filter=table_filter, match_overrides=match_overrides,
) )
return _runner
register_step( # 合并 step 注册表(title / format / 描述相对稳定,单独列)
step_id="std_ind_001", _COMBINED_STEP_META = {
title="IND-001 · 身份证号校验", "std_ind_001": {
description=( "parent_id": "ind_001",
"title": "IND-001 · 身份证号校验",
"description": (
"GB 11643-1999 身份证号全维度校验:4 子检查合一(a 格式 / b 校验位 / c 出生日期 / d 15 位老证)。" "GB 11643-1999 身份证号全维度校验:4 子检查合一(a 格式 / b 校验位 / c 出生日期 / d 15 位老证)。"
"同一字段同一值多子检查失败的会合并成 1 行展示,error 字段拼全部子错误信息。" "同一字段同一值多子检查失败的会合并成 1 行展示,error 字段拼全部子错误信息。"
), ),
requires_db=True, llm_mode="none", required=False, order=100, # 国标字段规范 group 起始 100;合并 step 置首位 "order": 100,
fn=_run_standards_ind_001, "purpose": (
detail=StepDetail(
purpose=(
"对身份证号字段跑 IND-001-a(格式)、IND-001-b(校验位)、" "对身份证号字段跑 IND-001-a(格式)、IND-001-b(校验位)、"
"IND-001-c(出生日期)、IND-001-d(15 位老证)4 个子检查," "IND-001-c(出生日期)、IND-001-d(15 位老证)4 个子检查,"
"合并输出到 1 个 tab,相同 (table, column, value) 的多子失败合并为 1 行," "合并输出到 1 个 tab,相同 (table, column, value) 的多子失败合并为 1 行,"
"error 字段拼接所有子错误信息。" "error 字段拼接所有子错误信息。"
), ),
target="字段名含 id_card / id_card_no / id_number / identity_card 或注释匹配的字段", "target": "字段名含 id_card / id_card_no / id_number / identity_card 或注释匹配的字段",
check=( "check": (
"a) 18 位 + 6+8+3+1 字符结构\n" "a) 18 位 + 6+8+3+1 字符结构\n"
"b) ISO 7064 MOD 11-2 校验位\n" "b) ISO 7064 MOD 11-2 校验位\n"
"c) value[6:14] 解析为合法日期(含闰年 2-29)\n" "c) value[6:14] 解析为合法日期(含闰年 2-29)\n"
"d) 15 位老证提示(已自 1999-07-01 停发)" "d) 15 位老证提示(已自 1999-07-01 停发)"
), ),
format="GB 11643-1999(合并)/ GB 11643-1989(d 子项)", "format": "GB 11643-1999(合并)/ GB 11643-1989(d 子项)",
},
"std_ind_002": {
"parent_id": "ind_002",
"title": "IND-002 · 统一社会信用代码校验",
"description": (
"GB 32100-2015 统一社会信用代码全维度校验:3 子检查合一(a 格式 / b 校验位 / c 老代码兼容转换)。"
"同一字段同一值多子检查失败的会合并成 1 行展示,error 字段拼全部子错误信息。"
),
"order": 101,
"purpose": (
"对统一社会信用代码字段跑 IND-002-a(18 位 + 登记管理部门 / 机构类别编码)、"
"IND-002-b(ISO 7064 MOD 31-3 校验位)、IND-002-c(老代码兼容转换)"
"3 个子检查,合并输出到 1 个 tab,相同 (table, column, value) 的多子失败合并为 1 行。"
),
"target": "字段名含 uscc / credit_code / social_credit 或注释匹配的字段",
"check": (
"a) 18 位 + 第 1 位登记管理部门类别 + 第 2 位机构类别\n"
"b) ISO 7064 MOD 31-3 校验位(字符集 0-9 + A-Z 共 31 字符)\n"
"c) 9 位的旧工商注册号 / 组织机构代码 → 转换到 18 位的兼容检测"
),
"format": "GB 32100-2015",
},
"std_ind_003": {
"parent_id": "ind_003",
"title": "IND-003 · 手机号校验",
"description": (
"工信部手机号全维度校验:2 子检查合一(a 格式 / b 号段)。"
"同一字段同一值多子检查失败的会合并成 1 行,error 字段拼全部子错误信息。"
),
"order": 102,
"purpose": (
"对手机号字段跑 IND-003-a(11 位 + 1[3-9] 开头格式)、IND-003-b(号段在工信部已放出号段内)"
"2 个子检查,合并输出到 1 个 tab。"
),
"target": "字段名含 mobile / phone / cell_phone / tel 或注释匹配的字段",
"check": (
"a) 11 位数字 + 1[3-9]\\d{9} 格式(不区分 +86 前缀)\n"
"b) 前 7 位号段在中国移动 / 联通 / 电信 / 广电 已发布号段表里"
),
"format": "工信部《电信网编号计划》",
},
"std_ind_004": {
"parent_id": "ind_004",
"title": "IND-004 · 行政区划代码校验",
"description": (
"GB/T 2260 行政区划代码全维度校验:2 子检查合一(a 6 位格式 / b 编码存在性)。"
"同一字段同一值多子检查失败的会合并成 1 行,error 字段拼全部子错误信息。"
),
"order": 103,
"purpose": (
"对行政区划代码字段跑 IND-004-a(6 位数字结构)+ IND-004-b(前 2 位 / 前 6 位"
"在 GB/T 2260 现行版本里可查到)2 个子检查,合并输出到 1 个 tab。"
),
"target": "字段名含 xzqh / region_code / area_code / district_code 或注释匹配的字段",
"check": (
"a) 6 位数字(省 2 + 市 2 + 县区 2)\n"
"b) 前 6 位在 GB/T 2260 现行版本县及县以上行政区划代码表里"
),
"format": "GB/T 2260",
},
}
for _step_id, _meta in _COMBINED_STEP_META.items():
register_step(
step_id=_step_id,
title=_meta["title"],
description=_meta["description"],
requires_db=True, llm_mode="none", required=False, order=_meta["order"],
fn=_run_standards_combined(_meta["parent_id"]),
detail=StepDetail(
purpose=_meta["purpose"],
target=_meta["target"],
check=_meta["check"],
format=_meta["format"],
),
)
# ── IND-005 改名(去掉 -a 后缀) ──────────────────────────────────
# 2026-08-13 起:唯一子检查 IND-005-a(固定电话格式)改名为 std_ind_005,
# UI 标题去掉 -a(与 IND-001/002/003/004 视觉一致)。
# step_id 由 indicator_step_id("IND-005-a") = "std_ind_005_a" 改为 "std_ind_005"。
def _run_standards_ind_005(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides): # noqa: ARG001
from .step_impl.step7_standards import run_step7_for_indicator
# 把它当作普通单 indicator 跑(IND-005-a 没合并到任何复合 step)
# section_key 强制改成 "standards_ind_005",与 step_id 视觉对齐
res = run_step7_for_indicator(
"IND-005-a", cfg, dict_data=dict_data, log=log,
table_filter=table_filter, match_override=match_overrides.get("std_ind_005") if match_overrides else None,
)
# 端到端重命名:section_key / step_id / tab.key / tab.title
res["section_key"] = "standards_ind_005"
proto = res.get("data", {}).get("_protocol") if isinstance(res.get("data"), dict) else None
if proto is not None:
proto["step_id"] = "std_ind_005"
for tab in (proto.get("tabs") or []):
if tab.get("key") == "standards_ind_005_a":
tab["key"] = "standards_ind_005"
if (tab.get("title") or "").startswith("IND-005-a"):
tab["title"] = "IND-005 · GB/T 15835 固定电话格式"
return res
register_step(
step_id="std_ind_005",
title="IND-005 · 固定电话格式",
description=(
"GB/T 15835 固定电话号码格式校验(11 位 / 3-4 位区号 + 7-8 位号码)。"
"原 std_ind_005_a 于 2026-08-13 去掉 -a 后缀,改为 std_ind_005。"
),
requires_db=True, llm_mode="none", required=False, order=104,
fn=_run_standards_ind_005,
detail=StepDetail(
purpose="对固定电话号码字段跑固定电话格式校验(区号 + 主号),违规值以 code 列展示。",
target="字段名含 tel / phone / fixed_phone / landline 或注释匹配的字段",
check=(
"a) 区号 3-4 位(少数 5 位)+ 空格可选 + 主号 7-8 位\n"
"b) 区号在 GB/T 15835 已发布的区号表里"
),
format="GB/T 15835",
), ),
) )
......
...@@ -1567,19 +1567,48 @@ def _empty_indicator_step_result(standard_id: str, reason: str = "") -> dict: ...@@ -1567,19 +1567,48 @@ def _empty_indicator_step_result(standard_id: str, reason: str = "") -> dict:
return {"section_key": section_key, "data": data} return {"section_key": section_key, "data": data}
# ── IND-001 合并校验 ────────────────────────────────────────── # ── 子检查合并校验(IND-001 / -002 / -003 / -004)─────────────────
# 把 IND-001-a / -b / -c / -d 四个子检查合并为一个 orchestrator step: # 把同一国标下的多个子检查合并为一个 orchestrator step:
# 1. UI 只展示 1 个 checkbox「IND-001 · 身份证号」 # 1. UI 只展示 1 个 checkbox(标题去掉 -a/-b/-c 后缀)
# 2. 后端跑全部 4 个子检查 # 2. 后端跑全部子检查
# 3. 同一 (table, column, value) 多子检查失败 → 合并成 1 行,error 字段拼所有子检查的错误 # 3. 同一 (table, column, value) 多子检查失败 → 合并成 1 行,
# error 字段拼所有子检查的错误
# 4. 单 section / 单 tab,结果表里 rule_type 列显示「IND-001-a, IND-001-b」 # 4. 单 section / 单 tab,结果表里 rule_type 列显示「IND-001-a, IND-001-b」
# #
# 设计动机:用户在「分析结果」里看身份证号字段时,4 个子检查的违规能聚合到一行, # 设计动机:用户在「分析结果」里看字段时,多个子检查的违规能聚合到一行,
# 一眼看清「这条数据有哪些子维度不合规」,不用上下翻 4 个 tab。 # 一眼看清「这条数据有哪些子维度不合规」,不用上下翻 N 个 tab。
# #
# 2026-08-12 新增,scope 限定 IND-001(其余 a/b 后缀 indicator 暂不动)。 # 2026-08-12 新增;2026-08-13 扩展到 IND-002 / -003 / -004。
# parent_id: e.g. "ind_001" —— 决定 step_id (std_ind_001) / section_key (standards_ind_001) / tab key
# sub_ids: e.g. ("IND-001-a", ..., "IND-001-d") —— 待跑子检查
# title: UI 标题("IND-001 · 身份证号校验")
# format: 标准来源("GB 11643-1999" / "GB 32100-2015" / 工信部 / "GB/T 2260")
_IND_COMBINED: dict[str, dict] = {
"ind_001": {
"sub_ids": ("IND-001-a", "IND-001-b", "IND-001-c", "IND-001-d"),
"title": "IND-001 · 身份证号校验",
"format": "GB 11643-1999(合并 a/b/c/d)/ GB 11643-1989(d 子项兼容)",
},
"ind_002": {
"sub_ids": ("IND-002-a", "IND-002-b", "IND-002-c"),
"title": "IND-002 · 统一社会信用代码校验",
"format": "GB 32100-2015(含 /-c 老代码兼容转换)",
},
"ind_003": {
"sub_ids": ("IND-003-a", "IND-003-b"),
"title": "IND-003 · 手机号校验",
"format": "工信部(《电信网编号计划》)",
},
"ind_004": {
"sub_ids": ("IND-004-a", "IND-004-b"),
"title": "IND-004 · 行政区划代码校验",
"format": "GB/T 2260",
},
}
_IND_001_SUB_IDS = ("IND-001-a", "IND-001-b", "IND-001-c", "IND-001-d") # 旧名保留以兼容 orchestrator 现有 import;新代码请用 _IND_COMBINED["ind_001"]["sub_ids"]
_IND_001_SUB_IDS = _IND_COMBINED["ind_001"]["sub_ids"]
def _merge_violations_by_tuple(violations: list[dict]) -> list[dict]: def _merge_violations_by_tuple(violations: list[dict]) -> list[dict]:
...@@ -1633,26 +1662,37 @@ def _merge_violations_by_tuple(violations: list[dict]) -> list[dict]: ...@@ -1633,26 +1662,37 @@ def _merge_violations_by_tuple(violations: list[dict]) -> list[dict]:
return out return out
def _wrap_combined_ind_001( def _wrap_combined_indicator(
parent_id: str,
combined: dict, combined: dict,
sub_instances: list[BaseStandard], sub_instances: list[BaseStandard],
) -> dict: ) -> dict:
"""包装 IND-001 合并结果 → 1 section / 1 tab。 """包装 (parent_id 合并的) 子检查结果 → 1 section / 1 tab。
Tab 协议与单 indicator 基本一致(沿用 _make_single_indicator_tab 的两表结构), Tab 协议与单 indicator 基本一致(沿用 _make_single_indicator_tab 的两表结构),
但 KPI 加一个「子检查通过数」、summary 显示 4 个子检查的 ID。 但 KPI 加一个「子检查通过数」、summary 显示所有子检查的 ID。
"""
sub_names = [f"{s.standard_id} {s.standard_name}" for s in sub_instances]
title = "IND-001 · 身份证号校验"
# violations_by_indicator 行加一个 `sub_id_short` 字段便于前端 tooltip 显示简短名 Args:
parent_id: e.g. "ind_001" —— 决定 step_id / section_key / tab key / KPI css_class
combined: 来自 _run_round1_one 累加的 {summary, violations_by_indicator, violations, ...}
sub_instances: 跑过的子 BaseStandard 实例列表
"""
cfg = _IND_COMBINED.get(parent_id)
if cfg is None:
raise ValueError(f"_wrap_combined_indicator: 未知 parent_id={parent_id}")
sub_ids = cfg["sub_ids"]
title = cfg["title"]
# 子检查 ID 的共享前缀,用于 sub_short 提取("IND-001-a" → "a")
sub_prefix = sub_ids[0].rsplit("-", 1)[0] + "-" if sub_ids else ""
# violations_by_indicator 行加一个 `sub_short` 字段便于前端 tooltip 显示简短名
rows = [] rows = []
for r in (combined.get("violations_by_indicator") or []): for r in (combined.get("violations_by_indicator") or []):
rr = dict(r) rr = dict(r)
ind = rr.get("indicator", "") or "" ind = rr.get("indicator", "") or ""
# "IND-001-a" → "a" # "IND-001-a" → "a"
if ind.startswith("IND-001-"): if sub_prefix and ind.startswith(sub_prefix):
rr["sub_short"] = ind[-1] rr["sub_short"] = ind[len(sub_prefix):]
rows.append(rr) rows.append(rr)
wrapped = dict(combined) wrapped = dict(combined)
...@@ -1665,21 +1705,21 @@ def _wrap_combined_ind_001( ...@@ -1665,21 +1705,21 @@ def _wrap_combined_ind_001(
) )
tab = TabProtocol( tab = TabProtocol(
key="standards_ind_001", key=f"standards_{parent_id}",
title=title, title=title,
kpis=( kpis=(
KpiSpec( KpiSpec(
css_class="kpi-ind-001", css_class=f"kpi-{parent_id}",
label="违规字段", label="违规字段",
value="this.summary.violations", value="this.summary.violations",
), ),
KpiSpec( KpiSpec(
css_class="kpi-ind-001-fields", css_class=f"kpi-{parent_id}-fields",
label="检查字段", label="检查字段",
value="this.summary.fields_checked", value="this.summary.fields_checked",
), ),
KpiSpec( KpiSpec(
css_class="kpi-ind-001-subs", css_class=f"kpi-{parent_id}-subs",
label="子检查命中", label="子检查命中",
value="this.summary.indicators_used", value="this.summary.indicators_used",
), ),
...@@ -1687,8 +1727,8 @@ def _wrap_combined_ind_001( ...@@ -1687,8 +1727,8 @@ def _wrap_combined_ind_001(
summary=SummarySpec( summary=SummarySpec(
type="warning", type="warning",
template=( template=(
"IND-001 校验:检查 {{fields}} 个字段," f"{parent_id.upper().replace('_', '-')} 校验:检查 {{{{fields}}}} 个字段,"
"发现 {{violations}} 条违规(同一字段同一值多子检查失败的已合并为 1 行)。" "发现 {{{{violations}}}} 条违规(同一字段同一值多子检查失败的已合并为 1 行)。"
), ),
vars={ vars={
"fields": "this.summary.fields_checked", "fields": "this.summary.fields_checked",
...@@ -1696,7 +1736,7 @@ def _wrap_combined_ind_001( ...@@ -1696,7 +1736,7 @@ def _wrap_combined_ind_001(
}, },
), ),
tables=( tables=(
# 表 1:按子 indicator 聚合(4 行 = a/b/c/d) # 表 1:按子 indicator 聚合
TableSpec( TableSpec(
title="子检查汇总", title="子检查汇总",
source="this.violations_by_indicator", source="this.violations_by_indicator",
...@@ -1762,19 +1802,21 @@ def _wrap_combined_ind_001( ...@@ -1762,19 +1802,21 @@ def _wrap_combined_ind_001(
), ),
), ),
) )
return inject_protocol(wrapped, step_id="std_ind_001", tabs=[tab]) return inject_protocol(wrapped, step_id=f"std_{parent_id}", tabs=[tab])
def run_step7_for_ind_001( def run_step7_combined(
parent_id: str,
cfg: DBConfig, cfg: DBConfig,
dict_data: dict, dict_data: dict,
log: Callable | None = None, log: Callable | None = None,
table_filter: set[str] | None = None, table_filter: set[str] | None = None,
match_overrides: dict[str, dict] | None = None, match_overrides: dict[str, dict] | None = None,
) -> dict: ) -> dict:
"""IND-001 合并校验入口:跑 a/b/c/d 4 个子检查,合并为 1 section / 1 tab。 """合并子检查的入口:跑 parent_id 下全部子检查,合并为 1 section / 1 tab。
Args: Args:
parent_id: e.g. "ind_001" —— 决定 _IND_COMBINED 选哪个子集
cfg / dict_data / log / table_filter: 同 run_step7_for_indicator cfg / dict_data / log / table_filter: 同 run_step7_for_indicator
match_overrides: 支持三种 key: match_overrides: 支持三种 key:
- "std_ind_001" 合并 step 的整体 override(推荐) - "std_ind_001" 合并 step 的整体 override(推荐)
...@@ -1784,30 +1826,35 @@ def run_step7_for_ind_001( ...@@ -1784,30 +1826,35 @@ def run_step7_for_ind_001(
Returns: Returns:
{"section_key": "standards_ind_001", "data": <merged _protocol>} {"section_key": "standards_ind_001", "data": <merged _protocol>}
""" """
section_key = "standards_ind_001" cfg_meta = _IND_COMBINED.get(parent_id)
if cfg_meta is None:
raise ValueError(f"run_step7_combined: 未知 parent_id={parent_id}(不在 _IND_COMBINED 里)")
sub_ids = cfg_meta["sub_ids"]
section_key = f"standards_{parent_id}"
merged_step_id = f"std_{parent_id}"
columns = dict_data.get("data_dictionary", []) columns = dict_data.get("data_dictionary", [])
if table_filter: if table_filter:
columns = [c for c in columns if c["table_name"] in table_filter] columns = [c for c in columns if c["table_name"] in table_filter]
# ── 实例化 4 个子 indicator(缺失一个整体跳过) ── # ── 实例化所有子 indicator(缺失一个整体跳过) ──
sub_instances: list[BaseStandard] = [] sub_instances: list[BaseStandard] = []
for sid in _IND_001_SUB_IDS: for sid in sub_ids:
cls = load_standard_class(sid) cls = load_standard_class(sid)
if cls is None: if cls is None:
if log: if log:
log("WARN", log("WARN",
f"IND-001 合并: 找不到子标准 {sid},跳过", f"{parent_id.upper().replace('_', '-')} 合并: 找不到子标准 {sid},跳过",
step="standards_ind_001") step=section_key)
continue continue
sub_instances.append(cls()) sub_instances.append(cls())
if not sub_instances: if not sub_instances:
# 兜底:返回空数据 + 单 tab # 兜底:返回空数据 + 单 tab
empty = _empty_indicator_data() empty = _empty_indicator_data()
empty["summary"]["indicators_applied"] = len(_IND_001_SUB_IDS) empty["summary"]["indicators_applied"] = len(sub_ids)
empty["note"] = "IND-001 子标准全部未找到" empty["note"] = f"{parent_id.upper().replace('_', '-')} 子标准全部未找到"
return {"section_key": section_key, "data": _wrap_combined_ind_001(empty, [])} return {"section_key": section_key, "data": _wrap_combined_indicator(parent_id, empty, [])}
# ── 合并桶 ── # ── 合并桶 ──
combined: dict = { combined: dict = {
...@@ -1824,7 +1871,7 @@ def run_step7_for_ind_001( ...@@ -1824,7 +1871,7 @@ def run_step7_for_ind_001(
if log: if log:
log("INFO", log("INFO",
f"IND-001 合并校验启动: 共 {len(sub_instances)} 个子检查 " f"{parent_id.upper().replace('_', '-')} 合并校验启动: 共 {len(sub_instances)} 个子检查 "
f"({', '.join(s.standard_id for s in sub_instances)})", f"({', '.join(s.standard_id for s in sub_instances)})",
step=section_key) step=section_key)
...@@ -1835,7 +1882,7 @@ def run_step7_for_ind_001( ...@@ -1835,7 +1882,7 @@ def run_step7_for_ind_001(
override = None override = None
if match_overrides: if match_overrides:
override = ( override = (
match_overrides.get("std_ind_001") match_overrides.get(merged_step_id)
or match_overrides.get(sid) or match_overrides.get(sid)
or match_overrides.get(indicator_step_id(sid)) or match_overrides.get(indicator_step_id(sid))
) )
...@@ -1860,12 +1907,27 @@ def run_step7_for_ind_001( ...@@ -1860,12 +1907,27 @@ def run_step7_for_ind_001(
if log: if log:
log("INFO", log("INFO",
f"IND-001 合并完成: {combined['summary']['violations']} 条违规, " f"{parent_id.upper().replace('_', '-')} 合并完成: {combined['summary']['violations']} 条违规, "
f"{combined['summary']['fields_checked']} 个字段, " f"{combined['summary']['fields_checked']} 个字段, "
f"{combined['summary']['indicators_used']} 个子检查命中", f"{combined['summary']['indicators_used']} 个子检查命中",
step=section_key) step=section_key)
return { return {
"section_key": section_key, "section_key": section_key,
"data": _wrap_combined_ind_001(combined, sub_instances), "data": _wrap_combined_indicator(parent_id, combined, sub_instances),
} }
# 保留旧名以兼容 orchestrator 现有调用;推荐新代码用 run_step7_combined
def run_step7_for_ind_001(
cfg: DBConfig,
dict_data: dict,
log: Callable | None = None,
table_filter: set[str] | None = None,
match_overrides: dict[str, dict] | None = None,
) -> dict:
"""IND-001 合并校验入口(保留为 run_step7_combined 的薄包装)。"""
return run_step7_combined(
"ind_001", cfg, dict_data,
log=log, table_filter=table_filter, match_overrides=match_overrides,
)
\ No newline at end of file
Markdown is supported
0%
or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment