Commit ba75ed1b authored by Data Governance Dev's avatar Data Governance Dev

feat(standards): IND-001 a/b/c/d 合并为单分析项目 + 4 子检查合一

需求:勾一个 IND-001 checkbox 跑全部 4 子检查;
结果合并展示,单条违规里显示「哪些子检查失败」。

改动:
- step7_standards.py: 新增 _IND_001_SUB_IDS / _merge_violations_by_tuple
  / _wrap_combined_ind_001 / run_step7_for_ind_001,单 (table,col,value)
  多子检查失败行合并,rule_type 列改为 'IND-001-a, IND-001-b',
  error 列改为 '[a] 长度 19 ≠ 18\n[b] 校验位错误' 拼接原文
- orchestrator.py: 注册 std_ind_001 step(order=100,llm=none),
  跳过 IND-001-a/b/c/d 自动注册,避免 4 个独立 section 还在
- analysis_tree.json: 国标字段规范组删 4 行 IND-001-a/b/c/d,加 1 行 std_ind_001
- routes.py: /api/match-config 过滤未注册为 step 的子 indicator;
  把 IND-001 a/b/c/d 的 applies_to_fields / comment_keywords 聚合到 std_ind_001
  的输入框,避免合并后前端输入框丢失
- 修复 double-wrap bug:_run_standards_ind_001 不再二次包裹
  run_step7_for_ind_001 原返回值(已含 {section_key, data})
- WORKLOG 追加 4 段:合并设计 / 关键词输入框回归 / 标题 '合并 a/b/c/d' 去除 /
  结果不显示修复(double-wrap 根因 + 烟测)
parent b968bb24
...@@ -48,3 +48,4 @@ LLM 不是「增强」,对部分步骤是核心依赖: ...@@ -48,3 +48,4 @@ LLM 不是「增强」,对部分步骤是核心依赖:
- 每做一个任务就在工作记录中记录一次 - 每做一个任务就在工作记录中记录一次
- 不要自动提交,我说提交再提交 - 不要自动提交,我说提交再提交
- 每次提交要写清除修改内容 - 每次提交要写清除修改内容
- 每次踩坑记录一下
...@@ -2,6 +2,198 @@ ...@@ -2,6 +2,198 @@
> 任务做完一次记一次。最近的在最上面。 > 任务做完一次记一次。最近的在最上面。
## 2026-08-12 · IND-001 合并后「结果又不显示」修复(double-wrap bug)
### 问题
合并 IND-001 后第一次跑:进度条显示 `已完成: 1/1`,但结果区一直 `暂无结果`,违规字段总数 = 0。
- 切换「-a / -b / -c / -d」4 个独立 step 还在 analysis_tree.json 时,至少能见到结果
- 合并后单独勾选 `std_ind_001`,section 进 sections 但前端读取不到任何 row
### 根因
[web/core/orchestrator.py:608](web/core/orchestrator.py#L608) 的合并 step 闭包错误地**双重包装**了返回值:
```python
def _run_standards_ind_001(...):
from .step_impl.step7_standards import run_step7_for_ind_001
data = run_step7_for_ind_001(cfg, dict_data=..., log=..., ...) # ← 已返回 {section_key, data}
return {"section_key": "standards_ind_001", "data": data} # ← 又包了一层
```
orchestrator 主循环 [orchestrator.py:881-883](web/core/orchestrator.py#L881-L883) 期望 step 返回 `{"section_key", "data"}`:
```python
section_key = output.get("section_key") # "standards_ind_001"
sections[section_key] = output["data"] # 这里 data 已是嵌套 dict,不再有 _protocol
```
→ `sections["standards_ind_001"]._protocol` 不存在 → 前端 `displayedTabs` 过滤掉 → `flatResultRows` 为空 → `暂无结果`。
对比 `_run_custom_value_check` ([orchestrator.py:592](web/core/orchestrator.py#L592)),被调函数返回的是纯 data(不带 `section_key`),所以那里包一层是对的;IND-001 的 `run_step7_for_ind_001` 自带 `section_key`,不能再包。
### 修复
[web/core/orchestrator.py:608-613](web/core/orchestrator.py#L608-L613) 去掉多余包装:
```python
def _run_standards_ind_001(...):
from .step_impl.step7_standards import run_step7_for_ind_001
# run_step7_for_ind_001 已返回 {section_key, data},直接透传(不要再 wrap)
return run_step7_for_ind_001(
cfg, dict_data=dict_data, log=log,
table_filter=table_filter, match_overrides=match_overrides,
)
```
### 烟测
空数据字典下跑 `run_step7_for_ind_001`,再走 `json.dumps(..., default=...)` 模拟 FastAPI 序列化:
```
json section_key: standards_ind_001
json protocol.step_id: std_ind_001
json tabs: 1
tab.key= standards_ind_001 tab.title= IND-001 · 身份证号校验
summary.violations: 0
```
`section_key` / `protocol.step_id` / `tabs[0].key` 三者对齐后,前端 `displayedTabs` → `flatResultRows` 即能正常渲染「国标字段规范」分组下的 IND-001 行。
---
## 2026-08-12 · IND-001 a/b/c/d 合并为单分析项目(4 子检查合一)
### 需求
> 用户反馈:「把分析指标的 IND-001 的 abcd 合成一个分析项目,但是都要做这些分析,在结果中显示哪里不合规(可能涉及修改公共结果 tab 的接口)」
原先 IND-001-a / -b / -c / -d 是 4 个独立 orchestrator step → 4 个独立 section / 4 个 tab:
- 选 1 个 checkbox 只能跑 1 个子检查(如选了 `-a` 就看不到校验位错误)
- 看一条身份证号数据违规,要切 4 个 tab 才看全
用户期望:勾一个 checkbox → 后端跑全部 4 子检查 → 结果合并展示,**单条违规明细里明确显示「哪些子检查失败」**。
### 设计决策(用户已确认)
1. **scope 限定 IND-001**:其余 a/b 后缀 indicator(IND-002/003/004/006/013/014/016)暂不动,保持独立 step
2. **同 (table, column, value) 多子检查失败 → 合并成 1 行**:
- `rule_type` 列改为 `"IND-001-a, IND-001-b"`(逗号分隔所有失败子检查 ID)
- `error` 列改为 `"[a] 长度 19 ≠ 18\n[b] 校验位错误,应为 X"`(拼接所有子错误,带子检查短标)
- 前端直接展示 error 文本,无需新协议
3. **不动 `RenderSpec.kind` 协议**:复用现有 `code` / `text` kind,避免破坏公共 tab 接口
4. **违规明细表保留 `rule_type` 列**:列名改为「失败子检查」,告知用户哪些子检查失败
### 修改 / 新增的 3 个文件
**[web/core/step_impl/step7_standards.py](web/core/step_impl/step7_standards.py)** — 文件末尾新增 IND-001 合并模块
```python
_IND_001_SUB_IDS = ("IND-001-a", "IND-001-b", "IND-001-c", "IND-001-d")
def _merge_violations_by_tuple(violations: list[dict]) -> list[dict]:
"""同 (table, column, value) 多子检查失败 → 1 行;rule_type / error 字段拼接。"""
# 按 (table_name, column_name, value) 三元组分组
# 单条 → 原样返回;多条 → 取首行为基底,拼接 rule_type + standard_name + error
# error 格式:「[a] 长度 19 ≠ 18\n[b] 校验位错误,应为 X」(带子 ID 短标)
# 额外字段 merged_sub_count 记录合并条数(前端可选展示,留作扩展点)
def _wrap_combined_ind_001(combined, sub_instances) -> dict:
"""1 section / 1 tab 协议包装。
KPIs: 违规字段 / 检查字段 / 子检查命中
summary: 「IND-001 合并校验(4 子检查合一):检查 X 字段,发现 Y 条违规」
tables: 子检查汇总(4 行 = a/b/c/d)+ 违规明细(rule_type 列名改「失败子检查」)
"""
def run_step7_for_ind_001(cfg, dict_data, log, table_filter, match_overrides) -> dict:
"""主入口。
- 逐子检查调 _run_round1_one(复用现有抽样/字段匹配逻辑)
- 收集所有 violations / violations_by_indicator
- _merge_violations_by_tuple 合并同 tuple 行
- match_overrides 三种 key 都接受:「std_ind_001」/「IND-001-a」/「std_ind_001_a」
- 返回 {section_key: "standards_ind_001", data: <wrapped>}
"""
```
**[web/core/orchestrator.py](web/core/orchestrator.py)** — 跳过子指标注册 + 注册复合 step
- `_register_indicator_steps()` 内新增 `_IND_001_SUBS = ("IND-001-a", ..., "-d")`,遇到即 `continue`(不计入 `same_group_count`)
- 新增 `_run_standards_ind_001(*, cfg, dict_data, ...)` 闭包 runner(捕获标准 + 透传 log/filter/overrides)
- 在 `_register_indicator_steps()` 调用之后,新增 `register_step(step_id="std_ind_001", order=100, ...)`
- `order=100`:国标字段规范 group 起始位置,置首位
- title: 「IND-001 · 身份证号(合并 a/b/c/d)」
- detail: 4 子检查的 purpose / target / check / format 串接
**[web/configs/analysis_tree.json](web/configs/analysis_tree.json)** — 4 → 1
- 国标字段规范 group 的 children:原 4 行(`std_ind_001_a/b/c/d`)替换为 1 行 `{ "step_id": "std_ind_001" }`
- 总 children 数 12 → 9
### 验证
```bash
$ python -c "from web.core.orchestrator import get_step_defs; print([s.step_id for s in get_step_defs() if '001' in s.step_id])"
['std_ind_001'] # 4 个子 step 全部不再注册;只有合并 step 在
```
```bash
$ python -c "from web.core.step_impl.step7_standards import _merge_violations_by_tuple as m; \
print(len(m([
{'table_name':'t1','column_name':'id_card','value':'X1','rule_type':'IND-001-a','error':'长度 19 ≠ 18'},
{'table_name':'t1','column_name':'id_card','value':'X1','rule_type':'IND-001-b','error':'校验位错'},
{'table_name':'t1','column_name':'id_card','value':'X2','rule_type':'IND-001-a','error':'长度 17'},
])))"
2 # 3 → 2(X1 两条合并为 1)
```
```bash
$ PYTHONIOENCODING=utf-8 python -c "
from web.core.step_impl.step7_standards import run_step7_for_ind_001
from web.core.db_adapter import DBConfig
r = run_step7_for_ind_001(DBConfig('mysql','h',1,'u','','d'), dict_data={}, log=None)
print(r['section_key'], r['data']['_protocol']['tabs'][0]['title'])"
standards_ind_001 IND-001 · 身份证号(合并 a/b/c/d)
```
### 端到端 tabKey 对齐
| 链路 | 字段 | 值 |
|---|---|---|
| UI 勾选 | step_id | `std_ind_001` |
| 协议 step_id | `_protocol.step_id` | `std_ind_001` |
| section_key | `sections` dict key | `standards_ind_001` |
| tab key | `_protocol.tabs[0].key` | `standards_ind_001` |
| displayedTabs | tabKey / stepId | `standards_ind_001` / `std_ind_001` |
| flatResultRows | 树叶子 row.tabKey | `standards_ind_001` |
`displayedTabs` 用 `proto.step_id ∈ planned` 过滤;用户勾了 `std_ind_001` → 该 section 入选。`flatResultRows` 用 `tabByStep[step_id]` 反查 → 找到 tab → 渲染左树叶子 + 右 pane。
### 已知约束 / 取舍
- **.gitignore**:`web/configs/db_defaults.yaml` 不入库(生产凭证);不影响本次改动
- **前端公共 tab 接口未扩展**:本任务靠 `rule_type` 列改为「失败子检查」+ `error` 拼接完成展示,**没有新增 `RenderSpec.kind`**。后续若要做更复杂的子状态可视化(如每子检查独立色 tag),可考虑新增 `sub_markers` kind
- **scope 不外扩**:IND-002/003/004/006/013/014/016 的 a/b/c 后缀仍走独立 step。模式已抽象成可复用 API(`_IND_001_SUB_IDS` / `_merge_violations_by_tuple` / `run_step7_for_ind_001`),后续按需扩展即可
### 后续微调(用户反馈)
> 「不显示后面的(abcd),这个概念就要去掉,一个分析项目就是一个项目」+「后面的关键词输入框怎么没了」
两处细节修正:
1. **去「合并 a/b/c/d」字样**(用户:「一个分析项目就是一个项目」)
- [web/core/orchestrator.py](web/core/orchestrator.py) `title`:「IND-001 · 身份证号(合并 a/b/c/d)」→「IND-001 · 身份证号校验」
- [web/core/step_impl/step7_standards.py](web/core/step_impl/step7_standards.py) tab title / summary template:去掉「(4 子检查合一)」/「合并校验」字样,summary 改为「IND-001 校验:检查 N 个字段,发现 N 条违规」
2. **恢复关键词输入框**(merge 之前 IND-001-a/b/c/d 各自有字段匹配输入框,merge 后只剩 1 个 step 但 UI 没显示输入框)
- 原因:`/api/match-config` 端点遍历 `_list_std_step_ids_safe()` 输出每个 BaseStandard 子类的输入框配置,IND-001-a/b/c/d 不再注册为独立 step 但仍出现在 YAML 配置里 → 端点不再生成对应条目 → UI 看不到输入框
- 修复:[web/api/routes.py](web/api/routes.py) `get_match_config()` 末尾新增「合并 step」块,把 IND-001-a/b/c/d 4 个 YAML 条目的 `applies_to_fields` 聚合到 `std_ind_001.names`(去重),让 UI 输入框有默认值
- 同时:循环里加 `if sid not in registered_step_ids: continue`,过滤掉未注册为 step 的孤儿条目(之前会留下 `std_ind_001_a/b/c/d` 4 个无主输入框)
后端 `[web/core/step_impl/step7_standards.py](web/core/step_impl/step7_standards.py)` `run_step7_for_ind_001` 已经支持 3 种 key 的 override lookup(`std_ind_001` / `IND-001-a` / `std_ind_001_a`),所以用户编辑 `std_ind_001` 输入框后 → 后端正确应用到 4 个子检查。
### 重新验证
```bash
$ python -c "from web.api.routes import get_match_config; import asyncio; r=asyncio.run(get_match_config()); print(r['steps']['std_ind_001'])"
{'names': 'id_card, id_card_no, id_number, identity_card, sfz_hm', 'comments': '', 'skip': False}
```
```bash
$ python -c "from web.core.orchestrator import _STEPS; print(_STEPS['std_ind_001'].title)"
IND-001 · 身份证号校验
```
确认 UI 现在会显示:「IND-001 · 身份证号校验」+ 字段匹配输入框(默认带 `id_card, id_card_no, id_number, identity_card, sfz_hm`)。
## 2026-08-12 · 自定义规则 v2:字段类型分类 + 递归条件树 ## 2026-08-12 · 自定义规则 v2:字段类型分类 + 递归条件树
### 需求 ### 需求
......
...@@ -216,9 +216,16 @@ async def get_match_config(): ...@@ -216,9 +216,16 @@ async def get_match_config():
# ── Step 7:每个 indicator 的 step_id(如 "std_ind_001_a")──── # ── Step 7:每个 indicator 的 step_id(如 "std_ind_001_a")────
# YAML 配置里的 key 是 standard_id(如 "IND-001-a"),需要 map 到 step_id。 # YAML 配置里的 key 是 standard_id(如 "IND-001-a"),需要 map 到 step_id。
# 约定:step_id = "std_ind_" + standard_id.lower().replace("-", "_") # 约定:step_id = "std_ind_" + standard_id.lower().replace("-", "_")
# 只输出「实际注册为 orchestrator step」的 indicator(IND-001-a/b/c/d
# 已合并到 std_ind_001,不再单独注册 —— 见下方「合并 step」块)
from ..core.orchestrator import get_step_defs as _get_step_defs
registered_step_ids = {s.step_id for s in _get_step_defs()}
for sid, meta in getattr(_list_std_step_ids_safe(), "__iter__", lambda: [])() or []: for sid, meta in getattr(_list_std_step_ids_safe(), "__iter__", lambda: [])() or []:
# meta: {"id": "IND-001-a", "name": "...", ...} # meta: {"id": "IND-001-a", "name": "...", ...}
std_id = meta.get("id", "") std_id = meta.get("id", "")
# 跳过未注册为 step 的子 indicator(合并 step 会单独处理)
if sid not in registered_step_ids:
continue
rec = step7.get(std_id) or {} rec = step7.get(std_id) or {}
if rec.get("_skip"): if rec.get("_skip"):
by_step[sid] = {"names": "", "comments": "", "skip": True} by_step[sid] = {"names": "", "comments": "", "skip": True}
...@@ -257,6 +264,30 @@ async def get_match_config(): ...@@ -257,6 +264,30 @@ async def get_match_config():
"skip": False, "skip": False,
} }
# ── 合并 step:std_ind_001(IND-001 a/b/c/d 聚合)───────────
# 4 个子 indicator 的 YAML 配置聚合到 1 个 step 输入框。
# 用户编辑的 override 在后端 run_step7_for_ind_001 里会下推到各子检查。
ind001_sub_ids = ("IND-001-a", "IND-001-b", "IND-001-c", "IND-001-d")
ind001_names: list[str] = []
ind001_comments: list[str] = []
ind001_seen_n: set = set()
ind001_seen_c: set = set()
for sub_id in ind001_sub_ids:
rec = step7.get(sub_id) or {}
for n in (rec.get("applies_to_fields") or []):
if n not in ind001_seen_n:
ind001_names.append(n)
ind001_seen_n.add(n)
for c in (rec.get("comment_keywords") or []):
if c not in ind001_seen_c:
ind001_comments.append(c)
ind001_seen_c.add(c)
by_step["std_ind_001"] = {
"names": ", ".join(ind001_names),
"comments": ", ".join(ind001_comments),
"skip": False,
}
# ── 其余 step(merge / empty / missing_comments 等)无需输入框 ── # ── 其余 step(merge / empty / missing_comments 等)无需输入框 ──
# 前端按 step_id 找不到时按 skip=True 处理即可 # 前端按 step_id 找不到时按 skip=True 处理即可
......
...@@ -18,10 +18,7 @@ ...@@ -18,10 +18,7 @@
"description": "IND-001 ~ IND-005 系列:按 GB / GA 标准做字段值级别合规校验(格式 / 校验位 / 出生日期 / 15 位兼容 / 号段 / 编码存在性 / 固定电话 / 老代码兼容)", "description": "IND-001 ~ IND-005 系列:按 GB / GA 标准做字段值级别合规校验(格式 / 校验位 / 出生日期 / 15 位兼容 / 号段 / 编码存在性 / 固定电话 / 老代码兼容)",
"default_expand": true, "default_expand": true,
"children": [ "children": [
{ "step_id": "std_ind_001_a" }, { "step_id": "std_ind_001" },
{ "step_id": "std_ind_001_b" },
{ "step_id": "std_ind_001_c" },
{ "step_id": "std_ind_001_d" },
{ "step_id": "std_ind_002_a" }, { "step_id": "std_ind_002_a" },
{ "step_id": "std_ind_002_b" }, { "step_id": "std_ind_002_b" },
{ "step_id": "std_ind_002_c" }, { "step_id": "std_ind_002_c" },
......
...@@ -540,8 +540,12 @@ def _register_indicator_steps() -> None: ...@@ -540,8 +540,12 @@ def _register_indicator_steps() -> None:
过滤: 过滤:
- 跳过 std_id 以 "STD-" 开头的已拆分旧插件(applies_to_fields=[]) - 跳过 std_id 以 "STD-" 开头的已拆分旧插件(applies_to_fields=[])
- 跳过 IND-001-a / -b / -c / -d(已被合并到 `std_ind_001` 复合 step,2026-08-12 起)
- 保留 IND-301 / IND-302 等无 applies_to_fields 的跨字段 indicator - 保留 IND-301 / IND-302 等无 applies_to_fields 的跨字段 indicator
""" """
# IND-001 的 4 个子指标合并为一个;其余 indicator 仍按 1 个 step / 1 个 tab 注册
_IND_001_SUBS = ("IND-001-a", "IND-001-b", "IND-001-c", "IND-001-d")
prev_group: str | None = None prev_group: str | None = None
same_group_count = 0 same_group_count = 0
base_order: int = 0 base_order: int = 0
...@@ -551,6 +555,12 @@ def _register_indicator_steps() -> None: ...@@ -551,6 +555,12 @@ def _register_indicator_steps() -> None:
if std_id.startswith("STD-"): if std_id.startswith("STD-"):
logger.info(f"跳过已 deprecated 旧插件: {std_id}(已被 IND-* 拆分)") logger.info(f"跳过已 deprecated 旧插件: {std_id}(已被 IND-* 拆分)")
continue continue
# 跳过 IND-001 子指标(已合并到 std_ind_001 复合 step)
if std_id in _IND_001_SUBS:
logger.info(
f"跳过子 indicator: {std_id}(已合并到 std_ind_001 复合 step)"
)
continue
group = meta.get("group", "通用") group = meta.get("group", "通用")
if group != prev_group: if group != prev_group:
same_group_count = 0 same_group_count = 0
...@@ -590,6 +600,48 @@ def _register_indicator_steps() -> None: ...@@ -590,6 +600,48 @@ def _register_indicator_steps() -> None:
_register_indicator_steps() _register_indicator_steps()
# ── IND-001 复合 step(合并 a/b/c/d 4 个子检查) ──────────────
# 2026-08-12 起:UI 上只展示 1 个 checkbox「IND-001 · 身份证号」,
# 后端内部跑全部 4 个子检查,结果按 (table, column, value) 合并展示,
# 同一个值若同时违反多个子检查 → 合并成 1 行,error 字段拼所有子错误。
# scope 限定 IND-001,其余 a/b 后缀 indicator 暂不动。
def _run_standards_ind_001(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides): # noqa: ARG001
from .step_impl.step7_standards import run_step7_for_ind_001
# run_step7_for_ind_001 已返回 {section_key, data},直接透传(不要再 wrap)
return run_step7_for_ind_001(
cfg, dict_data=dict_data, log=log,
table_filter=table_filter, match_overrides=match_overrides,
)
register_step(
step_id="std_ind_001",
title="IND-001 · 身份证号校验",
description=(
"GB 11643-1999 身份证号全维度校验:4 子检查合一(a 格式 / b 校验位 / c 出生日期 / d 15 位老证)。"
"同一字段同一值多子检查失败的会合并成 1 行展示,error 字段拼全部子错误信息。"
),
requires_db=True, llm_mode="none", required=False, order=100, # 国标字段规范 group 起始 100;合并 step 置首位
fn=_run_standards_ind_001,
detail=StepDetail(
purpose=(
"对身份证号字段跑 IND-001-a(格式)、IND-001-b(校验位)、"
"IND-001-c(出生日期)、IND-001-d(15 位老证)4 个子检查,"
"合并输出到 1 个 tab,相同 (table, column, value) 的多子失败合并为 1 行,"
"error 字段拼接所有子错误信息。"
),
target="字段名含 id_card / id_card_no / id_number / identity_card 或注释匹配的字段",
check=(
"a) 18 位 + 6+8+3+1 字符结构\n"
"b) ISO 7064 MOD 11-2 校验位\n"
"c) value[6:14] 解析为合法日期(含闰年 2-29)\n"
"d) 15 位老证提示(已自 1999-07-01 停发)"
),
format="GB 11643-1999(合并)/ GB 11643-1989(d 子项)",
),
)
register_step( register_step(
step_id="custom_value_check", step_id="custom_value_check",
title="自定义规则(字段值包含关键字)", title="自定义规则(字段值包含关键字)",
......
...@@ -1565,3 +1565,307 @@ def _empty_indicator_step_result(standard_id: str, reason: str = "") -> dict: ...@@ -1565,3 +1565,307 @@ def _empty_indicator_step_result(standard_id: str, reason: str = "") -> dict:
"tabs": [], "tabs": [],
} }
return {"section_key": section_key, "data": data} return {"section_key": section_key, "data": data}
# ── IND-001 合并校验 ──────────────────────────────────────────
# 把 IND-001-a / -b / -c / -d 四个子检查合并为一个 orchestrator step:
# 1. UI 只展示 1 个 checkbox「IND-001 · 身份证号」
# 2. 后端跑全部 4 个子检查
# 3. 同一 (table, column, value) 多子检查失败 → 合并成 1 行,error 字段拼所有子检查的错误
# 4. 单 section / 单 tab,结果表里 rule_type 列显示「IND-001-a, IND-001-b」
#
# 设计动机:用户在「分析结果」里看身份证号字段时,4 个子检查的违规能聚合到一行,
# 一眼看清「这条数据有哪些子维度不合规」,不用上下翻 4 个 tab。
#
# 2026-08-12 新增,scope 限定 IND-001(其余 a/b 后缀 indicator 暂不动)。
_IND_001_SUB_IDS = ("IND-001-a", "IND-001-b", "IND-001-c", "IND-001-d")
def _merge_violations_by_tuple(violations: list[dict]) -> list[dict]:
"""合并同一 (table_name, column_name, value) 多子检查失败行为 1 行。
每行最终结构:
- rule_type: "IND-001-a, IND-001-b" (所有失败子检查,逗号分隔)
- error: "[a] 长度 18 ≠ 18\\n[b] 校验位错误,应为 X" (所有子错误拼接)
- standard_name: 同上,逗号分隔
其他字段(table_name / column_name / value / column_type / column_comment / table_comment)
取第一条命中的值(同一 tuple 这些字段必然一致)。
单子检查失败的行保持原 schema 不变(rule_type 仍是单个 ID),前端无需感知差异。
"""
if not violations:
return []
groups: dict[tuple, list[dict]] = {}
for v in violations:
key = (
v.get("table_name", ""),
v.get("column_name", ""),
v.get("value", ""),
)
groups.setdefault(key, []).append(v)
out: list[dict] = []
for _key, rows in groups.items():
if len(rows) == 1:
# 单条:原样返回(去掉 _merge 后会临时加的字段,但单条路径本就不加,安全)
out.append(dict(rows[0]))
continue
# 多条:取首行作为基底,拼接 rule_type / error / standard_name
base = dict(rows[0])
rule_types = [r.get("rule_type", "") for r in rows]
errors = [r.get("error", "") for r in rows]
names = [r.get("standard_name", "") for r in rows]
# 保留稳定的子 ID 顺序(按 _IND_001_SUB_IDS 顺序),避免随机顺序影响 UI
# 如果顺序不重要就保留 dict 插入顺序
base["rule_type"] = ", ".join(t for t in rule_types if t)
base["standard_name"] = " | ".join(n for n in names if n)
base["error"] = "\n".join(
f"[{rt.split('-')[-1]}] {err}"
for rt, err in zip(rule_types, errors)
if rt and err
)
# 记录合并条数(前端可选展示,本次不用,留作扩展点)
base["merged_sub_count"] = len(rows)
out.append(base)
return out
def _wrap_combined_ind_001(
combined: dict,
sub_instances: list[BaseStandard],
) -> dict:
"""包装 IND-001 合并结果 → 1 section / 1 tab。
Tab 协议与单 indicator 基本一致(沿用 _make_single_indicator_tab 的两表结构),
但 KPI 加一个「子检查通过数」、summary 显示 4 个子检查的 ID。
"""
sub_names = [f"{s.standard_id} {s.standard_name}" for s in sub_instances]
title = "IND-001 · 身份证号校验"
# violations_by_indicator 行加一个 `sub_id_short` 字段便于前端 tooltip 显示简短名
rows = []
for r in (combined.get("violations_by_indicator") or []):
rr = dict(r)
ind = rr.get("indicator", "") or ""
# "IND-001-a" → "a"
if ind.startswith("IND-001-"):
rr["sub_short"] = ind[-1]
rows.append(rr)
wrapped = dict(combined)
wrapped["violations_by_indicator"] = rows
wrapped["summary"] = dict(combined.get("summary", {}))
# 加一个子检查通过数(命中 0 违规的 indicator 视为通过)
wrapped["summary"]["indicators_passed"] = (
wrapped["summary"].get("indicators_applied", 0)
- wrapped["summary"].get("indicators_used", 0)
)
tab = TabProtocol(
key="standards_ind_001",
title=title,
kpis=(
KpiSpec(
css_class="kpi-ind-001",
label="违规字段",
value="this.summary.violations",
),
KpiSpec(
css_class="kpi-ind-001-fields",
label="检查字段",
value="this.summary.fields_checked",
),
KpiSpec(
css_class="kpi-ind-001-subs",
label="子检查命中",
value="this.summary.indicators_used",
),
),
summary=SummarySpec(
type="warning",
template=(
"IND-001 校验:检查 {{fields}} 个字段,"
"发现 {{violations}} 条违规(同一字段同一值多子检查失败的已合并为 1 行)。"
),
vars={
"fields": "this.summary.fields_checked",
"violations": "this.summary.violations",
},
),
tables=(
# 表 1:按子 indicator 聚合(4 行 = a/b/c/d)
TableSpec(
title="子检查汇总",
source="this.violations_by_indicator",
default_sort={"prop": "violation_rate", "order": "descending"},
max_height=400,
columns=(
ColumnSpec(prop="indicator", label="子检查", width=140,
render=RenderSpec(kind="code")),
ColumnSpec(prop="indicator_name", label="名称",
min_width=200, show_overflow_tooltip=True),
ColumnSpec(prop="fields_count", label="检查字段数",
width=110, align="center"),
ColumnSpec(prop="violating_fields_count", label="违规字段数",
width=120, align="center"),
ColumnSpec(
label="违规率", prop="violation_rate", width=170,
align="center",
render=RenderSpec(
kind="percent",
value_path="violation_rate",
precision=2,
thresholds=(
Threshold(gt=0.5, color="#f56c6c"),
Threshold(gt=0.2, color="#e6a23c"),
Threshold(gt=0.0, color="#67c23a"),
),
),
),
),
),
# 表 2:违规明细(同一字段同一值多子检查失败 → 1 行,error 拼全部子错误)
TableSpec(
title="违规明细",
source="this.violations",
default_sort=None,
max_height=520,
columns=(
ColumnSpec(
label="表名 / 表注释", width=260,
render=RenderSpec(
kind="two_line",
main_path="table_name",
sub_path="table_comment",
sub_empty_text="(无表注释)",
),
),
ColumnSpec(
label="字段名 / 字段注释", width=240,
render=RenderSpec(
kind="two_line",
main_path="column_name",
sub_path="column_comment",
sub_empty_text="(无字段注释)",
),
),
ColumnSpec(prop="rule_type", label="失败子检查", width=180,
render=RenderSpec(kind="code")),
ColumnSpec(prop="value", label="违规值", width=180,
render=RenderSpec(kind="code")),
ColumnSpec(prop="error", label="违规原因(多子检查已拼接)",
min_width=280, show_overflow_tooltip=True),
),
),
),
)
return inject_protocol(wrapped, step_id="std_ind_001", tabs=[tab])
def run_step7_for_ind_001(
cfg: DBConfig,
dict_data: dict,
log: Callable | None = None,
table_filter: set[str] | None = None,
match_overrides: dict[str, dict] | None = None,
) -> dict:
"""IND-001 合并校验入口:跑 a/b/c/d 4 个子检查,合并为 1 section / 1 tab。
Args:
cfg / dict_data / log / table_filter: 同 run_step7_for_indicator
match_overrides: 支持三种 key:
- "std_ind_001" 合并 step 的整体 override(推荐)
- "IND-001-a" / "-b" / "-c" / "-d" 单子检查 override(向后兼容老 payload)
- "std_ind_001_a" 等 orchestrator 风格的 step_id key(向后兼容)
Returns:
{"section_key": "standards_ind_001", "data": <merged _protocol>}
"""
section_key = "standards_ind_001"
columns = dict_data.get("data_dictionary", [])
if table_filter:
columns = [c for c in columns if c["table_name"] in table_filter]
# ── 实例化 4 个子 indicator(缺失一个整体跳过) ──
sub_instances: list[BaseStandard] = []
for sid in _IND_001_SUB_IDS:
cls = load_standard_class(sid)
if cls is None:
if log:
log("WARN",
f"IND-001 合并: 找不到子标准 {sid},跳过",
step="standards_ind_001")
continue
sub_instances.append(cls())
if not sub_instances:
# 兜底:返回空数据 + 单 tab
empty = _empty_indicator_data()
empty["summary"]["indicators_applied"] = len(_IND_001_SUB_IDS)
empty["note"] = "IND-001 子标准全部未找到"
return {"section_key": section_key, "data": _wrap_combined_ind_001(empty, [])}
# ── 合并桶 ──
combined: dict = {
"summary": {
"indicators_applied": len(sub_instances),
"indicators_used": 0,
"fields_checked": 0,
"violations": 0,
},
"violations_by_indicator": [],
"violations": [],
"indicators_info": [],
}
if log:
log("INFO",
f"IND-001 合并校验启动: 共 {len(sub_instances)} 个子检查 "
f"({', '.join(s.standard_id for s in sub_instances)})",
step=section_key)
# ── 逐子检查跑(共用 _run_round1_one,避免重写抽样 SQL / 字段匹配) ──
for std in sub_instances:
sid = std.standard_id
# 3 种 key 都接受
override = None
if match_overrides:
override = (
match_overrides.get("std_ind_001")
or match_overrides.get(sid)
or match_overrides.get(indicator_step_id(sid))
)
single = _run_round1_one(
cfg, columns, std, log,
match_override=override,
)
combined["violations_by_indicator"].extend(single["violations_by_indicator"])
combined["violations"].extend(single["violations"])
combined["summary"]["fields_checked"] += single["summary"]["fields_checked"]
combined["indicators_info"].extend(single["indicators_info"])
# ── 合并同一 (table, column, value) 多子检查失败的行 ──
combined["violations"] = _merge_violations_by_tuple(combined["violations"])
combined["summary"]["indicators_used"] = sum(
1 for v in combined["violations_by_indicator"]
if (v.get("total_invalid") or 0) > 0
)
combined["summary"]["violations"] = len(combined["violations"])
combined["violations"] = combined["violations"][:500]
if log:
log("INFO",
f"IND-001 合并完成: {combined['summary']['violations']} 条违规, "
f"{combined['summary']['fields_checked']} 个字段, "
f"{combined['summary']['indicators_used']} 个子检查命中",
step=section_key)
return {
"section_key": section_key,
"data": _wrap_combined_ind_001(combined, sub_instances),
}
\ No newline at end of file
Markdown is supported
0%
or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment