Commit 53f6205d authored by Data Governance Dev's avatar Data Governance Dev

refactor(web): 「字段长度检查」按字段类型拆 7 个独立 step + 3 层嵌套

后端:
  - step6_length_check.py: LENGTH_RULES 扁平列表 → 7 个 _BUILTIN_CATEGORIES
    (len_id_card / len_uscc / len_mobile / len_xzqh / len_postal / len_email / len_bank_card)
    每个分类独立 rules + UI 元数据 + expected_length + standard
  - orchestrator.py: 单 length_check step → 7 个 len_* 独立 step
    (工厂函数 _run_length_check_category(cat_id));旧的 length_check 保留为 hidden
  - match_config.py: 新增 get_length_categories_rules(cat_id);旧 get_length_rules() 保留为兜底
  - standards_match.yaml: step6_length.rules → step6_length.categories.<cat_id>.rules
    每分类 1 个 name_pattern + 1 个 comment_keyword(区划留 2 个)
  - routes.py /api/analysis-tree: 加递归 3 层嵌套(group → subgroup → leaf)
    新增 is_subgroup: True 显式标记(避免靠 step_id 缺失推断)
    /api/match-config: 给每个 len_* 单独返回 names/comments 默认值

前端:
  - analysis_tree.json: 「字段长度检查」从 leaf 改为 subgroup
    (key=length_check_group, 含 7 个 len_* 子 step)
  - index.html: flatCheckRows + flatResultRows 支持 3 层(_level 0/1/2)
    toggleGroupCheck / setAllChecks / setAllGroupsExpanded 改为递归处理子组叶子
    flatResultRows fallback 分支抽出 _push_leaf_row(item, level) 助手
  - style.css: 加 .tree-row--subgroup / .subgroup-title / .subgroup-count / .subgroup-description

Bug fix:
  - step_id vs tab_key 混用导致 displayedTabs 过滤掉 7 个 length tab
    section_key='length_len_id_card' / tab_key='length_len_id_card' / step_id='len_id_card' 三者分离
    _wrap() 加 step_id 参数(默认沿用 tab_key,向后兼容旧 length_check)

WORKLOG 记录所有改动 + 5 个踩坑(subgroup 识别 / YAML 关键词精简 / fallback 分支复用 /
CSS .tree-row--subgroup / LengthRule.__dict__ 转 dict 抛错 / step_id 与 tab_key 混用)
parent 1bb3ca98
...@@ -2,6 +2,130 @@ ...@@ -2,6 +2,130 @@
> 任务做完一次记一次。最近的在最上面。 > 任务做完一次记一次。最近的在最上面。
## 2026-08-14 · 「字段长度检查」拆分为 7 个子分类(3 层嵌套)
> 用户反馈:原来「字段长度检查」一个 tab 把 身份证 / USCC / 手机 / 区划 等不同字段类型
> 混在一个表里看(检查对象和检测值混在一起),要求:
> 1. 在「字段长度检查」下加一级,按字段类型拆成独立子检查
> 2. 同样字段名 / 表注释目标,每个子检查独立可编辑
> 3. YAML 配置加默认关键词,但不要太多 —— 每个 1-2 个
### 改动总览
| 层 | 文件 | 变更 |
|---|---|---|
| 后端-规则 | [web/core/step_impl/step6_length_check.py](web/core/step_impl/step6_length_check.py) | LENGTH_RULES 扁平列表 → 7 个 `_BUILTIN_CATEGORIES`(每分类独立 rules + UI 元数据) |
| 后端-注册 | [web/core/orchestrator.py](web/core/orchestrator.py) | 单 `length_check` step → 7 个独立 step(`len_id_card` ... `len_bank_card`),工厂函数 `_run_length_check_category(cat_id)` |
| 配置 | [web/core/match_config.py](web/core/match_config.py) | 新增 `get_length_categories_rules(cat_id)`;旧 `get_length_rules()` 保留为兜底 |
| 配置 | [web/configs/standards_match.yaml](web/configs/standards_match.yaml) | `step6_length.rules` → `step6_length.categories.<cat_id>.rules`;每个分类 1 个 name_pattern + 1 个 comment_keyword |
| 路由 | [web/api/routes.py](web/api/routes.py) | `/api/analysis-tree` 加 3 层嵌套(递归 `_build_node` 处理 step_id 或 children);`/api/match-config` 给每个 len_* 单独返回 names/comments |
| 树配置 | [web/configs/analysis_tree.json](web/configs/analysis_tree.json) | 「字段长度检查」从 leaf 改为 subgroup(key=`length_check_group`,含 7 个子 step) |
| 前端 | [web/static/index.html](web/static/index.html) | `flatCheckRows` + `flatResultRows` 加 3 层(_level 0/1/2 + `is_subgroup`);`toggleGroupCheck/setAllChecks/setAllGroupsExpanded` 递归处理 |
| 样式 | [web/static/style.css](web/static/style.css) | 新增 `.tree-row--subgroup` / `.subgroup-title` / `.subgroup-count` / `.subgroup-description` |
### 7 个分类与标准
| 分类 step_id | 标题 | 标准 | 标准长度 | 默认关键词数 |
|---|---|---|---|---|
| `len_id_card` | 身份证号 | GB 11643-1999 | 18 | 1(身份证) |
| `len_uscc` | 统一社会信用代码 | GB 32100-2015 | 18 | 1(统一社会信用代码) |
| `len_mobile` | 手机号 | YD/T 1313 | 11 | 1(手机号) |
| `len_xzqh` | 行政区划代码 | GB/T 2260 | 6(默认);省 2 / 市 4 也覆盖 | 1(行政区划) |
| `len_postal` | 邮政编码 | GB/T 23703 | 6 | 1(邮政编码) |
| `len_email` | 电子邮件 | RFC 5321 | 50 | 1(邮箱) |
| `len_bank_card` | 银行卡号 | JR/T 0002 | 19 | 1(银行卡号) |
> 用户要求「每个给一两个默认关键词」—— YAML 里 1 个;区划 2 个(兼顾省级区划 / 区县区划)。
### 向后兼容
- 旧的 `length_check` step 仍注册为 `hidden=True`,UI 不显示,但代码保留「跑全部 7 个分类到一个 tab」的能力(用于回退或临时聚合查看)。
- `match_config.get_length_rules()`(旧扁平格式)仍保留为空返回;新代码用 `get_length_categories_rules(cat_id)`。
- `analysis_tree.json` 里 `length_check_group` 取代原 `length_check` leaf 位置(JSON 是配置源,不需要双写)。
### 验证
```bash
$ python -c "from web.core.orchestrator import get_step_defs; print([s.step_id for s in get_step_defs() if s.step_id.startswith('len_') or s.step_id=='length_check'])"
['len_id_card', 'len_uscc', 'len_mobile', 'len_xzqh', 'len_postal', 'len_email', 'len_bank_card', 'length_check'] # 8 个;length_check hidden
$ python -c "import json; print([g for g in json.load(open('web/configs/analysis_tree.json'))['groups'][0]['children']])"
[{'step_id': 'empty_fields'}, {'step_id': 'missing_comments'}, {'key': 'length_check_group', 'children': [7 个 len_*]}]
$ # /api/match-config 模拟
$ python -c "import asyncio; from web.api.routes import get_match_config; print(len(asyncio.run(get_match_config())['steps']))"
19 # 7 len_* + 9 std_ind_* + 3 其他(合并 step + IND-005/006/013)
```
### 踩坑
#### 1. `analysis-tree` 返回的字段命名混乱
最初想把 subgroup 的叶子直接放在 `children` 里,但前端 `_build_node` 函数区分「叶子(带 step_id)」vs「子组(带 children)」,需要一个明确标识。**新增 `is_subgroup: True` 字段**,递归时根据这个标志走不同分支,避免靠 `step_id` 缺失来推断(脆弱)。
#### 2. YAML 关键字数控制
最初在 YAML 给每个分类堆了 3-5 个关键词(如「身份证 / 公民身份 / 证件号码」),但用户明确说「不要太多,每个一两个」。**重新精简到 1 个**(区划留 2 个)。后续如果误命中,可以 UI 临时加 override,或改 YAML 加。
#### 3. flatResultRows 改写时漏改 fallback 分支
最初只改了主分支处理 subgroup,忘了 fallback「其他」组也走 `_push_leaf_row` 路径 —— 旧 `_level: 1` 写死为硬编码,且 `out.push({...})` 行内联展开,无法复用。新代码提了 `_push_leaf_row(item, level)` 助手函数,主分支 + fallback 分支共用,避免重复 4 段。
#### 4. CSS 编译时引用未定义
写完 `.subgroup-title` / `.subgroup-description` 时漏看 style.css 顶部 `@import` 链 —— 确认没遗漏后,才发现 style.css 没有 `.tree-row--subgroup`,**需要新加**(直接复用 `.tree-row` 的 flex 布局,单独覆盖 `background/font-weight`)。
#### 5. /api/match-config 兜底 LengthRule 对象 vs dict
旧代码用 `r.__dict__` 把 LengthRule dataclass 转 dict,但 LengthRule 的 `name_pattern` 是 `re.Pattern` 对象,dict 里有 `.pattern` 属性 —— 后面 `_extract_names_comments` 拿 `r.get("name_pattern")` 拿到的是 Pattern 对象,调 `.strip("^$")` 直接抛 AttributeError。
**修复**:兜底分支单独处理 LengthRule 对象,直接读 `rule.name_pattern.pattern` 字符串,避免走 dict 转换。
### 踩坑(重启后才发现):step_id 与 tab_key 混用 → displayedTabs 全部漏掉
用户重启服务后结果区仍然没有 7 个 length tab。日志显示后端完全跑通:
```
[INFO] 主步骤调度完成: 成功 [..., 'len_bank_card', 'len_email', 'len_id_card', 'len_mobile', 'len_postal', 'len_uscc', 'len_xzqh', ...]
[INFO] GET /api/jobs/.../result, sections=['length_len_id_card', 'length_len_uscc', ..., 'length_len_bank_card']
```
后端 section 里 7 个 length_* 全在,前端 displayedTabs 却把它们全过滤了。
**根因**:step6 的 `_wrap()` 同时把 `step_id` 和 `tab_key` 都设成 `f"length_{category_id}"`(`length_len_id_card` 这种),但前端 displayedTabs 用 `proto.step_id` 去 `planned.has()` 匹配:
```js
if (!planned.has(proto.step_id)) continue; // planned = {'len_id_card', ...}
```
`planned` 来自 orchestrator 的 `steps_planned`,里面是 `len_id_card`(短 ID),不是 `length_len_id_card`(带前缀)。所以 7 个 tab 全部被 `continue` 掉。
**修复**:把 `_wrap()` 拆成两个参数:
- `tab_key = "length_len_id_card"` —— section_key 与 UI tab 名(必须唯一)
- `step_id = "len_id_card"` —— 对齐 orchestrator 的 step_id,用于关联 `steps_planned`
```python
def _wrap(data, *, tab_key, tab_title, step_id=None):
if step_id is None: step_id = tab_key # 向后兼容(旧 length_check)
return inject_protocol(data, step_id=step_id, tabs=[_make_length_tab(key=tab_key, title=tab_title)])
```
跑 `run_step6_for_category` 时显式传 `step_id=category_id`。
**教训**:section_key / tab_key / step_id 是三件事:
- section_key = 后端 sections dict 的 key(路由用)
- tab_key = UI tab 的 key(DOM 渲染用)
- step_id = orchestrator 注册的 step 标识(与 steps_planned 对齐)
以前 length_check 单 step 时三者是同一个字符串(`length_issues`),没出问题;拆分 7 个分类后我图省事沿用同一字符串,把 step_id 也带上前缀,破坏了与 steps_planned 的对等关系。
**调试技巧**:怀疑过滤问题时,直接在日志里 print:
```python
for s in sections:
print(s, '->', s['_protocol'].get('step_id'))
print('planned =', steps_planned)
```
三秒定位问题,不用猜。
## 2026-08-14 · 异常值列「哪里错了」高亮 + tooltip 多行修复 ## 2026-08-14 · 异常值列「哪里错了」高亮 + tooltip 多行修复
> 用户截图反馈两步问题: > 用户截图反馈两步问题:
......
...@@ -127,6 +127,9 @@ async def list_steps(): ...@@ -127,6 +127,9 @@ async def list_steps():
# 分组拓扑从 web/configs/analysis_tree.json 读; # 分组拓扑从 web/configs/analysis_tree.json 读;
# 叶子节点的 title / llm_mode / required 从 orchestrator 已注册的 StepDef 解析。 # 叶子节点的 title / llm_mode / required 从 orchestrator 已注册的 StepDef 解析。
# 这样新增 / 改名 step 只动 orchestrator;改层级只动 JSON。 # 这样新增 / 改名 step 只动 orchestrator;改层级只动 JSON。
#
# 2026-08-14:支持 3 层嵌套(group → subgroup → leaf);
# 子组用 "key" + "children" 描述,没有 step_id;走父组的 default_expand。
@router.get("/analysis-tree") @router.get("/analysis-tree")
async def analysis_tree(): async def analysis_tree():
from ..core.orchestrator import get_step_defs from ..core.orchestrator import get_step_defs
...@@ -141,23 +144,9 @@ async def analysis_tree(): ...@@ -141,23 +144,9 @@ async def analysis_tree():
logger.warning(f"analysis_tree.json 解析失败: {e}") logger.warning(f"analysis_tree.json 解析失败: {e}")
return {"groups": [], "_warning": f"parse error: {e}"} return {"groups": [], "_warning": f"parse error: {e}"}
by_id = {sd.step_id: sd for sd in get_step_defs() if not sd.hidden} def _build_leaf(sd) -> dict:
referenced: set[str] = set()
out_groups: list[dict] = []
for g in raw.get("groups", []):
children: list[dict] = []
missing: list[str] = []
for node in g.get("children", []):
sid = node.get("step_id")
if not sid:
continue
sd = by_id.get(sid)
if not sd:
# JSON 引到了 hidden 的 step 或未注册的 step —— 不暴露给 UI
continue
referenced.add(sid)
detail = sd.detail detail = sd.detail
children.append({ return {
"step_id": sd.step_id, "step_id": sd.step_id,
"title": sd.title, "title": sd.title,
"description": sd.description, "description": sd.description,
...@@ -171,14 +160,49 @@ async def analysis_tree(): ...@@ -171,14 +160,49 @@ async def analysis_tree():
"format": detail.format} "format": detail.format}
if detail else None if detail else None
), ),
}) }
def _build_node(node: dict, referenced: set, by_id: dict) -> dict | None:
"""递归构造节点:叶子(带 step_id)或子组(带 children)。"""
sid = node.get("step_id")
if sid:
sd = by_id.get(sid)
if not sd:
# JSON 引到了 hidden 的 step 或未注册的 step —— 不暴露给 UI
return None
referenced.add(sid)
return _build_leaf(sd)
# 子组节点:必须含 children
children: list[dict] = []
for sub in (node.get("children") or []):
built = _build_node(sub, referenced, by_id)
if built is not None:
children.append(built)
return {
"key": node.get("key"),
"title": node.get("title"),
"description": node.get("description", ""),
"default_expand": bool(node.get("default_expand", False)),
"is_subgroup": True,
"checks": children,
}
by_id = {sd.step_id: sd for sd in get_step_defs() if not sd.hidden}
referenced: set[str] = set()
out_groups: list[dict] = []
for g in raw.get("groups", []):
children: list[dict] = []
for node in g.get("children", []):
built = _build_node(node, referenced, by_id)
if built is not None:
children.append(built)
out_groups.append({ out_groups.append({
"key": g.get("key"), "key": g.get("key"),
"title": g.get("title"), "title": g.get("title"),
"description": g.get("description", ""), "description": g.get("description", ""),
"default_expand": bool(g.get("default_expand", False)), "default_expand": bool(g.get("default_expand", False)),
"checks": children, "checks": children,
**({"missing_steps": missing} if missing else {}),
}) })
# 兜底:JSON 没引到的 step,自动追加到 "__未分组__" 组里,确保不丢 # 兜底:JSON 没引到的 step,自动追加到 "__未分组__" 组里,确保不丢
...@@ -189,19 +213,7 @@ async def analysis_tree(): ...@@ -189,19 +213,7 @@ async def analysis_tree():
"title": "未分组(JSON 未引到)", "title": "未分组(JSON 未引到)",
"description": "analysis_tree.json 没有覆盖到的 step;建议补到对应分组", "description": "analysis_tree.json 没有覆盖到的 step;建议补到对应分组",
"default_expand": True, "default_expand": True,
"checks": [ "checks": [_build_leaf(s) for s in sorted(orphan, key=lambda x: x.order)],
{
"step_id": s.step_id, "title": s.title, "description": s.description,
"llm_mode": s.llm_mode, "required": s.required, "order": s.order,
"detail": (
{"purpose": s.detail.purpose,
"target": s.detail.target,
"check": s.detail.check,
"format": s.detail.format}
if s.detail else None
),
} for s in sorted(orphan, key=lambda x: x.order)
],
}) })
return {"groups": out_groups} return {"groups": out_groups}
...@@ -223,6 +235,9 @@ async def get_match_config(): ...@@ -223,6 +235,9 @@ async def get_match_config():
cfg = _load() cfg = _load()
step7 = (cfg.get("step7_standards") or {}) if isinstance(cfg, dict) else {} step7 = (cfg.get("step7_standards") or {}) if isinstance(cfg, dict) else {}
# Step 6:2026-08-14 拆分类后,优先用新格式 categories.<cat_id>.rules;
# 旧格式 step6_length.rules(扁平)作为兜底(指向已被 hidden 的 length_check step)。
step6_cats = ((cfg.get("step6_length") or {}).get("categories") or {}) if isinstance(cfg, dict) else {}
step6_rules = ((cfg.get("step6_length") or {}).get("rules") or []) if isinstance(cfg, dict) else [] step6_rules = ((cfg.get("step6_length") or {}).get("rules") or []) if isinstance(cfg, dict) else []
# 把所有 step_id 收集起来:保证 UI 上每个 leaf 都有一份默认配置 # 把所有 step_id 收集起来:保证 UI 上每个 leaf 都有一份默认配置
...@@ -255,26 +270,68 @@ async def get_match_config(): ...@@ -255,26 +270,68 @@ async def get_match_config():
"skip": False, "skip": False,
} }
# ── Step 6:length_check 聚合所有规则的 names / comments ────── # ── Step 6(新格式):每个 length 分类单独返回 names / comments 默认值 ──────
# 从 step6_length.categories.<cat_id>.rules 抽 alternation 和 keywords。
# 兜底:如果新格式为空,回退到代码内置默认(list_categories)。
try:
from ..core.step_impl.step6_length_check import list_categories as _list_cats
_builtin_cats = _list_cats()
except Exception:
_builtin_cats = []
def _extract_names_comments(rules: list[dict]) -> tuple[list[str], list[str]]:
names: list[str] = []
comments: list[str] = []
for r in rules or []:
pat = r.get("name_pattern") or ""
body = pat.strip("^$")
for piece in body.split("|"):
p = piece.strip().rstrip("$").lstrip("^").strip("()")
if p and p not in names:
names.append(p)
for kw in (r.get("comment_keywords") or []):
if kw and kw not in comments:
comments.append(kw)
return names, comments
for cat in _builtin_cats:
cat_rules = ((step6_cats.get(cat.category_id) or {}).get("rules")) or []
if cat_rules:
# 新格式:YAML dict 列表
names, comments = _extract_names_comments(cat_rules)
else:
# 兜底:代码内置 LengthRule 对象列表
names = []
comments = []
for rule in cat.rules:
body = rule.name_pattern.pattern.strip("^$")
for piece in body.split("|"):
p = piece.strip().rstrip("$").lstrip("^").strip("()")
if p and p not in names:
names.append(p)
for kw in rule.comment_keywords:
if kw and kw not in comments:
comments.append(kw)
by_step[cat.category_id] = {
"names": ", ".join(names),
"comments": ", ".join(comments),
"skip": False,
}
# ── Step 6(兜底):旧的 length_check 聚合(hidden=True,UI 不显示但仍保留通道) ──
all_names: list[str] = [] all_names: list[str] = []
all_comments: list[str] = [] all_comments: list[str] = []
for r in step6_rules: for r in step6_rules:
# 从 regex pattern 抽出可读的 alternation(仅做显示,不强求一致)
# 例:^(id_?card|id_?number)$ → "id_card, id_number"
# 简单实现:拿 | 分隔的 alternation 串
pat = r.get("name_pattern") or "" pat = r.get("name_pattern") or ""
body = pat.strip("^$") body = pat.strip("^$")
for piece in body.split("|"): for piece in body.split("|"):
p = piece.strip().rstrip("$").lstrip("^") p = piece.strip().rstrip("$").lstrip("^").strip("()")
# 剥掉分组括号((?:...)/(...) 等)→ UI 显示干净 if p and p not in all_names:
p = p.strip("()")
# 把正则简化为可读示例(id_?card → id_card / id_card;这里只列原文)
if p:
all_names.append(p) all_names.append(p)
for kw in (r.get("comment_keywords") or []): for kw in (r.get("comment_keywords") or []):
if kw not in all_comments: if kw and kw not in all_comments:
all_comments.append(kw) all_comments.append(kw)
if all_names or all_comments:
by_step["length_check"] = { by_step["length_check"] = {
"names": ", ".join(all_names), "names": ", ".join(all_names),
"comments": ", ".join(all_comments), "comments": ", ".join(all_comments),
......
{ {
"_comment": "检查项分组配置。叶子节点用 step_id 引用 orchestrator 已注册的 step。改分组只动这个文件;新增 step 只动 orchestrator。两个源头互相解耦。", "_comment": "检查项分组配置。叶子节点用 step_id 引用 orchestrator 已注册的 step。改分组只动这个文件;新增 step 只动 orchestrator。两个源头互相解耦。2026-08-14 增加子组(subgroups)层级:字段长度检查拆为 7 个独立 step,放进 subgroups[].children。",
"groups": [ "groups": [
{ {
"key": "basic", "key": "basic",
...@@ -9,7 +9,21 @@ ...@@ -9,7 +9,21 @@
"children": [ "children": [
{ "step_id": "empty_fields" }, { "step_id": "empty_fields" },
{ "step_id": "missing_comments" }, { "step_id": "missing_comments" },
{ "step_id": "length_check" } {
"key": "length_check_group",
"title": "字段长度检查",
"description": "按字段类型拆分(身份证 / USCC / 手机 / 区划 / 邮编 / 邮箱 / 银行卡),单独勾选 + 单独配置(names / comments)。",
"default_expand": true,
"children": [
{ "step_id": "len_id_card" },
{ "step_id": "len_uscc" },
{ "step_id": "len_mobile" },
{ "step_id": "len_xzqh" },
{ "step_id": "len_postal" },
{ "step_id": "len_email" },
{ "step_id": "len_bank_card" }
]
}
] ]
}, },
{ {
......
...@@ -429,80 +429,78 @@ step7_standards: ...@@ -429,80 +429,78 @@ step7_standards:
IND-902: { _skip: true } IND-902: { _skip: true }
# ── Step 6:字段长度检查 ─────────────────────────────── # ── Step 6:字段长度检查(2026-08-14 按分类拆分)──────────────────
# 每个分类是一个独立 step(len_id_card / len_uscc / ...),
# 用户单独勾选 + 单独 override(互不污染)。
#
# 字段名走正则(精确的 alternation),注释走 substring。 # 字段名走正则(精确的 alternation),注释走 substring。
# 用户 UI override:names 会作为额外 substring 规则追加(每条 input 都触发, # 默认每个分类只列 1-2 个注释关键字(避免误命中);如需扩展,用户在 UI 加。
# expected_length 默认 18 —— 用户改实际长度时同步把 YAML 改对); # 用户 UI override:
# comments 作为额外 comment_keywords 追加到所有现有规则。 # - names → 追加为整条 substring 规则(expected_length = 分类默认值)
# - comments → 追加到本分类所有规则的 comment_keywords
#
# 旧的「step6_length.rules 扁平列表」仍兼容(get_length_rules() 走 fallback 路径);
# 优先用新格式。
step6_length: step6_length:
categories:
# ── 1. 身份证号 ── GB 11643-1999
len_id_card:
rules: rules:
- name_pattern: "^(id_?card|id_?number|identity_?card)$" - name_pattern: "^(id_?card|id_?number|identity_?card)$"
comment_keywords: [身份证, 公民身份]
expected_length: 18
standard: "GB 11643-1999"
description: "身份证号 18 位"
- name_pattern: "^(id_?card_?no|_?sfz_?hm)$"
comment_keywords: [身份证] comment_keywords: [身份证]
expected_length: 18 expected_length: 18
standard: "GB 11643-1999" standard: "GB 11643-1999"
description: "身份证号 18 位" description: "身份证号 18 位"
# ── 2. 统一社会信用代码 ── GB 32100-2015
len_uscc:
rules:
- name_pattern: "^(uscc|credit_?code|social_?credit_?code)$" - name_pattern: "^(uscc|credit_?code|social_?credit_?code)$"
comment_keywords: [统一社会信用代码, 社会信用代码, 信用代码] comment_keywords: [统一社会信用代码]
expected_length: 18 expected_length: 18
standard: "GB 32100-2015" standard: "GB 32100-2015"
description: "统一社会信用代码 18 位" description: "统一社会信用代码 18 位"
# ── 3. 手机号 ── YD/T 1313
len_mobile:
rules:
- name_pattern: "^mobile(_?phone)?$" - name_pattern: "^mobile(_?phone)?$"
comment_keywords: [手机号, 手机号码, 移动电话] comment_keywords: [手机号]
expected_length: 11
standard: "YD/T 1313"
description: "手机号 11 位"
- name_pattern: "^phone(_?no)?$"
comment_keywords: [联系电话]
expected_length: 11
standard: "YD/T 1313"
description: "手机号 11 位"
- name_pattern: "^tel(ephone)?$"
comment_keywords: [电话]
expected_length: 11 expected_length: 11
standard: "YD/T 1313" standard: "YD/T 1313"
description: "手机号 11 位" description: "手机号 11 位"
# ── 4. 行政区划代码 ── GB/T 2260
len_xzqh:
rules:
- name_pattern: "^(xzqhbm|adcode|district_?code)$" - name_pattern: "^(xzqhbm|adcode|district_?code)$"
comment_keywords: [行政区划, 行政区划代码, 地区编码] comment_keywords: [行政区划]
expected_length: 6 expected_length: 6
standard: "GB/T 2260" standard: "GB/T 2260"
description: "行政区划代码 6 位" description: "行政区划代码 6 位"
- name_pattern: "^province_?code$"
comment_keywords: [省级代码, 省代码, 省份编码]
expected_length: 2
standard: "GB/T 2260"
description: "省级代码 2 位"
- name_pattern: "^city_?code$"
comment_keywords: [市级代码, 市代码, 城市编码]
expected_length: 4
standard: "GB/T 2260"
description: "市级代码 4 位"
- name_pattern: "^region_?code$" - name_pattern: "^region_?code$"
comment_keywords: [区县级代码, 区县代码, 区县编码] comment_keywords: [区县代码]
expected_length: 6 expected_length: 6
standard: "GB/T 2260" standard: "GB/T 2260"
description: "区县级代码 6 位" description: "区县级代码 6 位"
# ── 5. 邮政编码 ── GB/T 23703
len_postal:
rules:
- name_pattern: "^zip_?code$" - name_pattern: "^zip_?code$"
comment_keywords: [邮政编码, 邮编] comment_keywords: [邮政编码]
expected_length: 6
standard: "GB/T 23703"
description: "邮政编码 6 位"
- name_pattern: "^post_?code$"
comment_keywords: [邮政编码, 邮编]
expected_length: 6 expected_length: 6
standard: "GB/T 23703" standard: "GB/T 23703"
description: "邮政编码 6 位" description: "邮政编码 6 位"
# ── 6. 电子邮件 ── RFC 5321
len_email:
rules:
- name_pattern: "^email$" - name_pattern: "^email$"
comment_keywords: [邮箱, 电子邮件, e-mail] comment_keywords: [邮箱]
expected_length: 50 expected_length: 50
standard: "RFC 5321" standard: "RFC 5321"
description: "电子邮件 ≤254 位,常用 ≤50" description: "电子邮件 ≤254 位,常用 ≤50"
# ── 7. 银行卡号 ── JR/T 0002
len_bank_card:
rules:
- name_pattern: "^bank_?card(_?no)?$" - name_pattern: "^bank_?card(_?no)?$"
comment_keywords: [银行卡号, 银行卡] comment_keywords: [银行卡号]
expected_length: 19 expected_length: 19
standard: "JR/T 0002" standard: "JR/T 0002"
description: "银行卡号 ≤19 位" description: "银行卡号 ≤19 位"
...@@ -108,7 +108,7 @@ def get_indicator_match(standard_id: str) -> dict[str, list[str]]: ...@@ -108,7 +108,7 @@ def get_indicator_match(standard_id: str) -> dict[str, list[str]]:
# ── Step 6 查询接口 ──────────────────────────────────────────── # ── Step 6 查询接口 ────────────────────────────────────────────
def get_length_rules() -> list[dict] | None: def get_length_rules() -> list[dict] | None:
"""获取 step6 的 LENGTH_RULES 配置。 """获取 step6 的 LENGTH_RULES 配置(旧的扁平格式,向后兼容)。
返回 list[dict](每条含 name_pattern / comment_keywords / expected_length / 返回 list[dict](每条含 name_pattern / comment_keywords / expected_length /
standard / description),配置文件缺失或 step6_length.rules 为空时返回 None standard / description),配置文件缺失或 step6_length.rules 为空时返回 None
...@@ -121,6 +121,32 @@ def get_length_rules() -> list[dict] | None: ...@@ -121,6 +121,32 @@ def get_length_rules() -> list[dict] | None:
return rules return rules
def get_length_categories_rules(category_id: str) -> list[dict] | None:
"""获取 step6 某个分类下的规则列表(新格式,2026-08-14 拆分后使用)。
配置格式:
step6_length:
categories:
len_id_card:
rules:
- name_pattern: "..."
comment_keywords: [...]
expected_length: 18
standard: "GB 11643-1999"
description: "..."
返回 list[dict],缺失或为空时返回 None → 调用方回退到代码内置默认。
"""
cfg = _load()
cats = ((cfg.get("step6_length") or {}).get("categories")) or {}
if not isinstance(cats, dict):
return None
rules = (cats.get(category_id) or {}).get("rules")
if not rules:
return None
return list(rules)
# ── 用户 override 合并 ──────────────────────────────────────── # ── 用户 override 合并 ────────────────────────────────────────
def merge_user_override( def merge_user_override(
base_atf: list[str] | None, base_atf: list[str] | None,
......
...@@ -125,9 +125,29 @@ def _run_missing_comments(*, cfg, dict_data, llm, log, cancel_event, table_filte ...@@ -125,9 +125,29 @@ def _run_missing_comments(*, cfg, dict_data, llm, log, cancel_event, table_filte
return {"section_key": "missing_comments", "data": data} return {"section_key": "missing_comments", "data": data}
def _run_length_check(*, cfg, dict_data, llm, log, cancel_event, table_filter, def _run_length_check_category(category_id: str):
match_overrides, sample_limit=None): # noqa: ARG001 """工厂:返回单个 length 分类的 runner(捕获 category_id 闭包)。
"""字段长度检查:纯字段定义匹配,不采样行数据,sample_limit 签名占位。"""
2026-08-14 拆分:把原本单一的 length_check step 拆成 7 个分类 step
(身份证 / USCC / 手机 / 区划 / 邮编 / 邮箱 / 银行卡),
每个分类单独注册、可单独勾选、有独立的 (names, comments) override。
"""
def _runner(*, cfg, dict_data, llm, log, cancel_event, table_filter, # noqa: ARG001
match_overrides, sample_limit=None):
from .step_impl.step6_length_check import run_step6_for_category
override = (match_overrides or {}).get(category_id)
out = run_step6_for_category(category_id, dict_data, log=log,
table_filter=table_filter,
match_override=override)
return out
return _runner
# 向后兼容:旧的 length_check step 仍注册但 hidden=True(不暴露 UI);
# 当前端 JSON 移除引用后可以彻底删。现在保留便于回退。
def _run_length_check(*, cfg, dict_data, llm, log, cancel_event, table_filter, # noqa: ARG001
match_overrides, sample_limit=None):
"""聚合 7 个分类到一个 tab —— 旧 UI 的兼容入口(hidden 后不再被 orchestrator 调用)。"""
from .step_impl.step6_length_check import run_step6 from .step_impl.step6_length_check import run_step6
override = (match_overrides or {}).get("length_check") override = (match_overrides or {}).get("length_check")
data = run_step6(dict_data, log=log, table_filter=table_filter, data = run_step6(dict_data, log=log, table_filter=table_filter,
...@@ -247,22 +267,56 @@ register_step( ...@@ -247,22 +267,56 @@ register_step(
format="纯规则,无 LLM;不依赖 llm.yaml api_key;可空列把 YES/Y → NULL(warning)、NO/N → NOT NULL(success)", format="纯规则,无 LLM;不依赖 llm.yaml api_key;可空列把 YES/Y → NULL(warning)、NO/N → NOT NULL(success)",
), ),
) )
register_step( # 字段长度检查:按字段类型分类,每个分类单独注册 step(2026-08-14 拆分)
step_id="length_check", # - 用户可单独勾选("我只关心身份证" → 只跑 len_id_card)
title="字段长度检查", # - 每个分类独立的 (names, comments) override(互不污染)
description="离线:识别字段定义长度超出国家标准(纯规则匹配,不依赖 LLM)", # - 每个分类独立 tab 渲染(避免 4 种字段混在一个表里)
requires_db=False, llm_mode="none", required=False, order=4, # order 4.x 让它们排在 missing_comments 之后
fn=_run_length_check, _LENGTH_CATEGORY_ORDER_BASE = 4
try:
from .step_impl.step6_length_check import list_categories as _list_length_cats
_LENGTH_CATEGORIES = _list_length_cats()
except Exception as _e: # pragma: no cover
logger.warning("加载 length 分类列表失败: %s", _e)
_LENGTH_CATEGORIES = []
for _i, _cat in enumerate(_LENGTH_CATEGORIES):
register_step(
step_id=_cat.category_id,
title=f"字段长度 · {_cat.title}",
description=_cat.description,
requires_db=False, llm_mode="none", required=False,
order=_LENGTH_CATEGORY_ORDER_BASE * 10 + _i,
fn=_run_length_check_category(_cat.category_id),
detail=StepDetail( detail=StepDetail(
purpose="识别字段定义长度超出国家标准的字段(如身份证 18 位但定义 VARCHAR(50) → 浪费空间 / 暗示长度不固定;或定义 VARCHAR(10) 但标准要求 18 → 必然写入失败),辅助 schema 整改。", purpose=f"识别【{_cat.title}】字段定义长度超出 {_cat.standard_short} 的字段"
f"(如定义 VARCHAR(50) 但 {_cat.standard_short} 规定 {_cat.expected_length} 位 → 浪费空间 / 暗示长度不固定;"
f"或定义 VARCHAR(10) 但标准要求 {_cat.expected_length} → 必然写入失败),辅助 schema 整改。",
target="information_schema.columns 全部字段定义(CHAR / VARCHAR / NVARCHAR 的 character_maximum_length)", target="information_schema.columns 全部字段定义(CHAR / VARCHAR / NVARCHAR 的 character_maximum_length)",
check=( check=(
"1. 字段名匹配已知国标字段名(身份证 / USCC / 手机号 / 区划 ...)\n" f"1. 字段名正则匹配(精确 alternation)\n"
"2. 与内置标准长度对比:身份证 18 / USCC 18 / 手机号 11 / 区划 6 / 婚姻代码 2\n" f"2. 字段注释关键字 substring 匹配(默认 1-2 个,用户可加)\n"
"3. 输出三类:定义过短(必失败)/ 准标(无变化)/ 定义过长(建议缩短)\n" f"3. 命中后与 {_cat.standard_short} 标准长度 {_cat.expected_length} 对比\n"
"4. 不修改 schema,仅产出整改建议清单" f"4. 输出两类:定义过长(建议缩短)/ 定义过短(必失败)\n"
f"5. 不修改 schema,仅产出整改建议清单"
),
format=f"对照 {_cat.standard_short}({_cat.title} {_cat.expected_length} 位);纯字符串匹配,无算法",
), ),
format="对照内置国标规则库(身份证 18 / USCC 18 / 手机号 11 / 区划 6 / 婚姻 2);纯字符串匹配,无算法", )
# 旧的聚合 length_check 仍注册(hidden=True,UI 不显示,但保留向后兼容通道)
register_step(
step_id="length_check",
title="字段长度检查(聚合,已弃用)",
description="聚合 7 个分类到单 tab;2026-08-14 起改用按分类 step",
requires_db=False, llm_mode="none", required=False, order=999,
fn=_run_length_check,
hidden=True,
detail=StepDetail(
purpose="聚合 7 个 length 分类 step 的输出到 1 个 tab —— 已弃用",
target="(同 len_*)",
check="(同 len_*)",
format="(同 len_*)",
), ),
) )
......
This diff is collapsed.
This diff is collapsed.
...@@ -228,6 +228,12 @@ body { ...@@ -228,6 +228,12 @@ body {
.tree-row--group { .tree-row--group {
font-size: 14px; font-size: 14px;
} }
.tree-row--subgroup {
font-size: 13px;
background: #fafbfc;
font-weight: 500;
color: #606266;
}
.tree-row--leaf { .tree-row--leaf {
font-size: 13px; font-size: 13px;
color: #303133; color: #303133;
...@@ -275,7 +281,12 @@ body { ...@@ -275,7 +281,12 @@ body {
font-weight: 600; font-weight: 600;
color: #303133; color: #303133;
} }
.group-count { .subgroup-title {
font-weight: 500;
color: #606266;
}
.group-count,
.subgroup-count {
margin-left: 8px; margin-left: 8px;
font-weight: normal; font-weight: normal;
color: #909399; color: #909399;
...@@ -283,6 +294,18 @@ body { ...@@ -283,6 +294,18 @@ body {
font-variant-numeric: tabular-nums; font-variant-numeric: tabular-nums;
flex-shrink: 0; flex-shrink: 0;
} }
.subgroup-description {
margin-left: 8px;
color: #909399;
font-size: 12px;
font-weight: normal;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
min-width: 0;
flex: 0 1 auto;
max-width: 560px;
}
.group-description { .group-description {
margin-left: 8px; margin-left: 8px;
color: #909399; color: #909399;
......
Markdown is supported
0%
or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment