Commit e7fb2a8c authored by Data Governance Dev's avatar Data Governance Dev

refactor(web): 收敛分析项 UI 至 3 组(hide/merge/rename 6 轮)

UI 从 7 组 / 30+ step 收敛到 3 组 / 15 step;隐藏 27 个 indicator。

合并 / 改名为独立 IND 号(移入 standards 组):
- IND-401 / IND-402 → std_ind_401(姓名校验,2 子检查合一)
- IND-006-b → std_ind_006(邮政编码)
- IND-013-a / -b / -c → std_ind_017 / -018 / -019
  (单位类别 GB/T 12402 / 经济类型 国统字〔2011〕86 号 / 行业代码 GB/T 4754)

隐藏 27 个 indicator(_HIDDEN_SUBS,暂离 UI,按需恢复):
- IND-016-a / -101 / -201..-204 / -301 / -302 / -501 / -502
- IND-601..-603 / -901 / -902
- IND-006-a / -008-a..-011-a(其他业务字段整组下架)
- IND-013-a..-c(已改名)
- IND-012-a / -014-a / -014-b / -015-a(单位法人经办人整组下架)

整组下架(5 个 group):
- 单位对公账户 字段规范
- 民政信息
- 证件信息
- 其他业务字段
- 单位 / 法人 / 经办人 字段规范

新增 / 重命名 group:
- 业务字段规范(含 std_ind_401 姓名 + std_ind_007_a 公积金账号)

后端改动:
- 抽 _run_standards_ind_renamed(source_id, target_step_id, target_title) 工厂
  合并 IND-005/006/013 的改名逻辑
- _run_standards_combined factory 处理 std_ind_401 合并子检查
- _IND_COMBINED['ind_401'] 配置(IND-401/IND-402 字符集 + 长度合一)
- _MERGED_SUBS / _HIDDEN_SUBS / mapped_step 同步更新

前端 / 配置:
- analysis_tree.json: 7 组 → 3 组;description / default_expand 同步精简
- routes.py /api/match-config: _COMBINED_GROUPS 聚合 + IND-005/006/013 改名映射

docs(WORKLOG): 6 轮完整记录 + 踩坑(log 闭包 bug 修复、sub_prefix 对无后缀 ID 的适配等)

UI 顺序:基础检查 → 国标字段规范(10 step)→ 业务字段规范(2 step)
parent a72a2ec1
...@@ -2,6 +2,518 @@ ...@@ -2,6 +2,518 @@
> 任务做完一次记一次。最近的在最上面。 > 任务做完一次记一次。最近的在最上面。
## 2026-08-13 · 「单位 / 法人 / 经办人 字段规范」整组下架
### 需求
> 用户原始反馈:「单位法人整个分析对象也隐藏」
### 设计决策
1. **IND-012-a / -014-a / -014-b / -015-a 临时隐藏**:放进 `_HIDDEN_SUBS`
2. **`business_unit` 整组删除**:4 个 child 全隐藏后组空了,与「单位对公账户」「民政信息」「证件信息」「其他业务字段」同样处理
### 改动
**[web/core/orchestrator.py](web/core/orchestrator.py)** — `_HIDDEN_SUBS` 追加 4 项
```python
_HIDDEN_SUBS = (
...,
"IND-012-a", # 单位名称
"IND-014-a", # 设立日期
"IND-014-b", # 启缴年月
"IND-015-a", # 发薪日
)
```
**[web/configs/analysis_tree.json](web/configs/analysis_tree.json)** — 删除整个 `business_unit` 组块
### 烟测
```
=== groups in file ===
basic | 基础检查 | 3
standards | 国标字段规范 | 10
business_main | 业务字段规范 | 2
← business_unit 组已删除
=== visible (hidden filter applied) ===
orphan: [] ← 无泄漏
=== hidden indicator check(27 个全检)===
IND-006-a / -008-a / -009-a / -010-a / -011-a /
IND-012-a / -014-a / -014-b / -015-a / -016-a /
IND-101 / -201 / -202 / -203 / -204 /
IND-301 / -302 / -501 / -502 /
IND-601 / -602 / -603 / -901 / -902 /
IND-013-a / -013-b / -013-c
→ 全部 OK hidden
```
### 后续
- `_HIDDEN_SUBS` 现含 **27 个 indicator**
- UI 上只剩 **3 组**:基础检查 / 国标字段规范 / 业务字段规范(核心)
- `analysis_tree.json` 现在只有 17 行左右(精简度大幅提升)
---
## 2026-08-13 · IND-013 三子项改名 IND-017/018/019 移入国标 + 「其他业务字段」整组下架
### 需求
> 用户原始反馈:「把 IND-013 的三个都放到国标里面,不要后面的 -a, -b, -c,直接增加 IND 号,但是不要和别的冲突了。然后其他业务字段整个隐藏」
### 设计决策
1. **IND-013 三子项各自分到独立 IND 号**:
- IND-013-a(GB/T 12402 单位类别)→ `std_ind_017`
- IND-013-b(国统字〔2011〕86 号 经济类型)→ `std_ind_018`
- IND-013-c(GB/T 4754 行业代码)→ `std_ind_019`
- 选 017/018/019 的理由:现有 IND 号 001-016 / 401 已被占用,017/018/019 是无冲突的连续号
- 复用现有 rename 模式(与 IND-005-a → std_ind_005 / IND-006-b → std_ind_006 同款):section_key / step_id / tab.key / tab.title 4 处同步改
2. **抽出 runner 工厂 `_run_standards_ind_renamed(source_id, target_step_id, target_title)`**:3 个改名 step 复用同一份改名逻辑(之前 IND-005/006 各写一份,这次抽象出来)
3. **「其他业务字段」整组下架**:5 个子项(006-a / 008-a / 009-a / 010-a / 011-a)全部隐藏 → 组空了,整组删除
- 与之前的「单位对公账户」「民政信息」「证件信息」组同款处理
### 改动
**[web/core/orchestrator.py](web/core/orchestrator.py)** — 4 处改动
1. `_MERGED_SUBS` 加 `IND-013-a` / `IND-013-b` / `IND-013-c` → 映射到 `std_ind_017` / `std_ind_018` / `std_ind_019`
2. `_HIDDEN_SUBS` 加 `IND-006-a` / `IND-008-a` / `IND-009-a` / `IND-010-a` / `IND-011-a`(共 20 个 indicator)
4. `mapped_step` dict 加 3 个映射
5. 抽 `_run_standards_ind_renamed(source_id, target_step_id, target_title)` runner 工厂
7. `register_step` 3 个新 block:std_ind_017(order=106)/ std_ind_018(order=107)/ std_ind_019(order=108)
**[web/configs/analysis_tree.json](web/configs/analysis_tree.json)** — 3 处改动
1. `standards` 组:加 017/018/019,描述更新为「IND-001 ~ IND-019 系列」
2. `business_unit` 组:移除 `std_ind_013_a/b/c`,描述更新为「IND-012 / IND-014 / IND-015 系列」
3. 删除整个 `business` 组(其他业务字段)
**[web/api/routes.py](web/api/routes.py)** — 加 IND-013-a/b/c → std_ind_017/018/019 聚合块(for 循环 3 次)
- 前端「字段匹配」输入框会聚合:
- std_ind_017: `unit_type / corp_type / company_type / enterprise_type / dwlx / dwlb / qylx` + `单位类型 / 单位类别 / ...`
- std_ind_018: `economic_type / econ_type / jjlx / jjxz` + `经济类型 / 经济性质 / ...`
- std_ind_019: `industry_code / industry / trade_code / hy_code / hybm / industry_category` + `行业代码 / 行业类别 / ...`
### 烟测
```
=== groups in file ===
basic | 基础检查 | 3
standards | 国标字段规范 | 10 ← + 017/018/019
business_main | 业务字段规范 | 2
business_unit | 单位 / 法人 / 经办人 字段规范 | 4 ← -013_a/b/c
← business 组已删除(其他业务字段整组下架)
=== visible (hidden filter applied) ===
orphan: [] ← 无泄漏
=== std_ind_017/018/019 rename check ===
std_ind_017 registered | order=106 | IND-017 · GB/T 12402 单位类别 值域
std_ind_018 registered | order=107 | IND-018 · 国统字〔2011〕86 号 经济类型 值域
std_ind_019 registered | order=108 | IND-019 · GB/T 4754 行业代码 格式
=== old std_ind_013_a/b/c ===
std_ind_013_a / _b / _c OK hidden(已合并改名)
=== hidden indicator check(20 个全检)===
IND-006-a / -008-a / -009-a / -010-a / -011-a /
IND-016-a / -101 / -201 / -202 / -203 / -204 /
IND-301 / -302 / -501 / -502 / -601 / -602 / -603 / -901 / -902 /
IND-013-a / -013-b / -013-c
→ 全部 OK hidden
```
### 踩坑 / 注意
- **`_run_standards_ind_renamed` 工厂初版错把 `log=None` 钉在闭包里**:第一次写时图省事把 log 当 closure 参数钉死,结果 dispatcher 传的 log 被丢掉(warning/error 日志全没了)。改成 `_runner(*, cfg, dict_data, llm, log, ...)` 接收 dispatcher 的 log,与 IND-005/006 模式一致
- **tab.key 推导从 `IND-013-a.lower().replace('-', '_') = "ind_013_a"` → `"standards_ind_013_a"`**:实际 `_wrap_single_indicator` 生成的是 `standards_ind_013_a`(不是 `standards_ind_013`),所以 prefix 匹配要包含 `_a` 后缀。代码里用 `startswith("standards_ind_013_a")` 形式 OK;判断时我用 `source_id.lower().replace('-', '_')` = `"ind_013_a"` 拼成 `"standards_ind_013_a"` 也对
- **business_unit 描述精简**:从「IND-012 ~ IND-015 系列:单位基本信息(名称 / 类型 / 经济类型 / 行业代码 / 设立日期 / 启缴年月 / 发薪日)」→「IND-012 / IND-014 / IND-015 系列:单位基本信息(名称 / 设立日期 / 启缴年月 / 发薪日)」
- **`_GROUP_ORDER_START["通用"] = 500` 现在只剩 5 个指标**(业务字段规范 / 单位法人经办人还在用);其他 group base 都变成 dead key,留着不影响恢复
### 后续
- `_HIDDEN_SUBS` 现含 **20 个 indicator**(含 IND-013-a/b/c + IND-006-a/008-a/009-a/010-a/011-a + 之前的)
- 国标字段规范组(standards)连续 10 个 IND 号:001~006 + 016-b + 017/018/019
- UI 上只剩 **4 组**:基础 / 国标 / 业务字段规范(核心) / 单位法人经办人
---
## 2026-08-13 · 新增「业务字段规范」组(401 + 007_a) + 旧业务组改名为「其他业务字段」
### 需求
> 用户原始反馈:「在国际标准规范下面增加'业务字段规范',把姓名校验(401)和公积金账号格式(007-a)放进去」
### 设计决策
1. **新增「业务字段规范」组(key=`business_main`)**,放在「国标字段规范」之后:
- 包含 `std_ind_401`(姓名校验)+ `std_ind_007_a`(公积金账号格式)
2. **旧「业务字段规范」组(key=`business`)改名为「其他业务字段」**:
- 移除 `std_ind_007_a`(已移到新组)
- 描述里去掉「个人公积金账号」字眼
- key 保留为 `business`(仅 title 改),历史引用无影响
3. **删除「身份信息」组**:`std_ind_401` 移走后该组空了,与「单位对公账户」「民政信息」「证件信息」同样处理
> ⚠ 新组与旧组标题相同会导致 UI 歧义。用户确认方案:旧组改名为「其他业务字段」
### 改动
**[web/configs/analysis_tree.json](web/configs/analysis_tree.json)** — 单文件改动,3 处
```json
// 1) 新增「业务字段规范」组(紧跟国标字段规范之后)
{
"key": "business_main",
"title": "业务字段规范",
"description": "IND-401 / IND-007-a:核心业务字段(姓名 / 个人公积金账号)",
"default_expand": true,
"children": [
{ "step_id": "std_ind_401" },
{ "step_id": "std_ind_007_a" }
]
}
// 2) 旧「业务字段规范」改名「其他业务字段」,移除 std_ind_007_a
{
"key": "business",
"title": "其他业务字段",
"description": "IND-006-a / IND-008-a ~ IND-011-a:其他业务字段(通讯地址 / 首次参加工作年月 / 用工类型 / 本地户籍 / 职工状态)",
"default_expand": true,
"children": [
{ "step_id": "std_ind_006_a" },
{ "step_id": "std_ind_008_a" },
{ "step_id": "std_ind_009_a" },
{ "step_id": "std_ind_010_a" },
{ "step_id": "std_ind_011_a" }
]
}
// 3) 删除「身份信息」组(原仅 std_ind_401,现已移走)
```
### 烟测
```
=== groups in file ===
basic | 基础检查 | 3
standards | 国标字段规范 | 7
business_main | 业务字段规范 (NEW) | 2 ← [std_ind_401, std_ind_007_a]
business | 其他业务字段 (RENAME) | 5 ← [006_a, 008_a..011_a]
business_unit | 单位 / 法人 / 经办人 字段规范 | 7
← identity 组已删除
=== visible (hidden filter applied) ===
orphan: [] ← 无泄漏
=== std_ind_401 / std_ind_007_a placement check ===
std_ind_401 -> business_main ✓
std_ind_007_a -> business_main ✓
```
### 后续
- 旧 `business` key 保留:未做依赖方迁移(前端的 analysis-tree 树完全 key 不敏感;后端 /api/analysis-tree 只读 JSON 不引用 key)
- 未来恢复「身份信息」组:把 `std_ind_401` 从 `business_main.children` 移走 + 新增 `identity` 组
- 未来合并回旧 key 命名:把 `business_main` 重命名为 `business`、旧 `business` 改成 `business_other` 之类 → 与本次对称操作
---
## 2026-08-13 · 证件信息整组下架 + IND-006-b 改名 IND-006 移入国标字段规范
### 需求
> 用户原始反馈:「证件信息整个不显示,把 006-b 移动到国标规范下面,编号改为 IND-006,不要后缀 -b」
### 设计决策
1. **IND-101 / -201 / -202 / -203 / -204 临时隐藏**:放进 `_HIDDEN_SUBS`
2. **`id_info` 整组删除**:5 个 child 全隐藏后组空了,跟「单位对公账户」「民政信息」同样处理
3. **IND-006-b 改名 + 移组**:
- step_id 由 `std_ind_006_b` 改为 `std_ind_006`(与 IND-005-a → IND-005 同款改名模式)
- 从「业务字段规范」移到「国标字段规范」,order = 105(接 IND-005 之后)
- IND-006-a(通讯地址)保留在「业务字段规范」不变 —— 用户只动了 b
- 4 处同步重命名:section_key / step_id / tab.key / tab.title
### 改动
**[web/core/orchestrator.py](web/core/orchestrator.py)** — 3 处改动
1. `_MERGED_SUBS` 加 `IND-006-b`
2. `_HIDDEN_SUBS` 加 `IND-101 / IND-201 / IND-202 / IND-203 / IND-204`
3. `mapped_step` dict 加 `"IND-006-b": "std_ind_006"`
4. 新增 `_run_standards_ind_006` runner + `register_step("std_ind_006", ...)`
- 镜像 IND-005 的模式:调 `run_step7_for_indicator("IND-006-b", ...)` → 端到端改名 section_key / step_id / tab.key / tab.title
**[web/configs/analysis_tree.json](web/configs/analysis_tree.json)** — 3 处改动
1. **`standards` 组**:新增 `std_ind_006`,描述里加「邮编」字眼
2. **`id_info` 组**:整组删除
3. **`business` 组**:移除 `std_ind_006_b`,描述里去掉「邮编」字眼(已上移到 standards)
**[web/api/routes.py](web/api/routes.py)** — 加 IND-006-b → std_ind_006 聚合块
- 镜像 IND-005 的处理:从 YAML key `IND-006-b` 拉 `applies_to_fields` / `comment_keywords` → 写入 `by_step["std_ind_006"]`
- 前端「字段匹配」输入框会自动显示 `postal_code, postcode, zip, zip_code, zipcode` + `邮编 / 邮政编码 / 单位邮编`
### 烟测
```
=== groups in file ===
basic | 基础检查 | 3
standards | 国标字段规范 | 7 (新增 std_ind_006)
identity | 身份信息 | 1
business | 业务字段规范 | 6 (去掉 std_ind_006_b)
business_unit | 单位 / 法人 / 经办人 字段规范 | 7
← id_info 组已删除
=== visible (hidden filter applied) ===
standards | [std_ind_001 .. -005, std_ind_006 (NEW), std_ind_016_b]
business | [std_ind_006_a, -007_a .. -011_a] ← 没了 006_b
orphan: []
=== registered std_ind_006 改名检查 ===
std_ind_006 registered: True
std_ind_006_b registered: False
=== hidden indicator check(15 个全检)===
IND-016-a / -101 / -201 / -202 / -203 / -204 /
-301 / -302 / -501 / -502 / -601 / -602 / -603 / -901 / -902
→ 全部 hidden OK,无 leak
=== match-config aggregation ===
IND-006-b YAML → std_ind_006
names=['postal_code','postcode','zip','zip_code','zipcode']
comments=['邮编','邮政编码','单位邮编']
```
### 踩坑 / 注意
- **`_derive_target` / `_derive_check` 里的 IND-101 / -201..-204 分支**:现在都是 dead code,留着不影响(恢复时还能用)
- **业务字段规范的 `purpose` 描述改了**:「通讯地址 / 邮编 / ...」→「通讯地址 / ...」(邮编移走了),JSON 里同步去掉了「邮编」字眼
- **`_GROUP_ORDER_START["证件信息"] = 200` 现在是 dead key**:留着不影响恢复
- **业务字段规范 child 顺序**:原顺序是 `006_a, 006_b, 007_a, ...`;现在变成 `006_a, 007_a, ...`(006_b 抽走了),UI 列表里 007_a 直接接 006_a
- **`_register_indicator_steps` 里的 STD-* skip 注释** 没改(仍提到「IND-001/002/003/004 各 a/b/c/d 子指标已合并」,描述准确,可保留)
### 后续
- `_HIDDEN_SUBS` 现含 **15 个 indicator**(含 5 个证件信息 + 民政 3 个 + 之前的 7 个)
- 合并/改名 step:`std_ind_001 / -002 / -003 / -004 / -005 / -006 / -401` 共 7 个
- UI 上只剩 **5 组**:基础 / 国标 / 身份 / 业务 / 单位法人经办人
---
## 2026-08-13 · 民政信息整组下架(IND-601 / -602 / -603 隐藏)
### 需求
> 用户原始反馈:「民政信息都隐藏 601,602,603」
### 设计决策
1. **IND-601 / -602 / -603 临时隐藏**:放进 `_HIDDEN_SUBS`,与 IND-016-a / -301 / -302 / -501 / -502 / -901 / -902 同集合
2. **`civil` 整组删除**:3 个 child 全部隐藏后组空了,跟「单位对公账户」组同样的处理(之前 IND-016/901/902 那次的同款决策)
3. **「民政信息」描述不保留**:组都没了,描述也跟着没了
### 改动
**[web/core/orchestrator.py](web/core/orchestrator.py)** — `_HIDDEN_SUBS` 追加 3 项
```python
_HIDDEN_SUBS = (
"IND-016-a",
"IND-301", "IND-302",
"IND-501", "IND-502",
"IND-601", "IND-602", "IND-603", # 新增
"IND-901", "IND-902",
)
```
**[web/configs/analysis_tree.json](web/configs/analysis_tree.json)** — 删除整个 `civil` 组块
- 与上一次「单位对公账户」组的处理一致:整组空 → 直接整组下架
### 烟测
```
=== groups in file ===
basic | 基础检查 | 3
standards | 国标字段规范 | 6
id_info | 证件信息 | 5
identity | 身份信息 | 1 (只有 std_ind_401)
business | 业务字段规范 | 7
business_unit | 单位 / 法人 / 经办人 字段规范 | 7
← civil 组已删除
=== visible (hidden filter applied) ===
(无 civil 组)
orphan: [] ← 无泄漏
=== hidden indicator check(10 个全检)===
IND-016-a / IND-301 / IND-302 / IND-501 / IND-502 /
IND-601 / IND-602 / IND-603 / IND-901 / IND-902
→ 全部 in _IND_COMBINED=False registered_as_step=False
```
### 后续
- `_HIDDEN_SUBS` 现含 10 个 indicator
- 「民政信息」组位置:JSON 顺序里 `civil` 组删除后,`business` 组紧跟在 `identity` 后面(视觉顺序:基础 → 国标 → 证件 → 身份 → 业务 → 单位法人经办人)
- 恢复路径:从 `_HIDDEN_SUBS` 移除 + 在 JSON 里恢复 `civil` 组(之前我给的身份信息组 / 单位对公账户组的恢复路径都是同一套)
---
## 2026-08-13 · IND-301 / -302 / -501 / -502 隐藏 + IND-401 / -402 合并为 std_ind_401
### 需求
> 用户原始反馈:「把 301,302,501,502 隐藏,401,402 合并」
### 设计决策
1. **IND-301 / -302 / -501 / -502 临时隐藏**:放进 `_HIDDEN_SUBS`(与 IND-016-a / 901 / 902 同集合,按需恢复)
2. **IND-401 + IND-402 合并为单 step `std_ind_401`**:
- 复用现有的 `_IND_COMBINED` + `_COMBINED_STEP_META` + `_run_standards_combined` 工厂
- 与 IND-001 ~ IND-004 合并结构对齐:UI 只 1 个 checkbox、1 个 section / 1 个 tab、同一字段同一值多子检查合并为 1 行
- **特殊点**:这一对没有 `-a/-b` 后缀,是直接以「姓名」「姓名长度」为 ID 的 2 个独立 indicator。`sub_prefix` 推导仍走得通:`"IND-401".rsplit("-", 1)[0] + "-"` = `"IND-"`,所以 `sub_short` 会显示成 `401` / `402`(之前 a/b/c 是字母);合并后 `error` 字段前缀变成 `[401] ...\n[402] ...`
3. **`identity` 组** 从 6 个 child 缩到 1 个;描述精简为「IND-401 · 姓名校验(字符集 + 长度,2 子检查合一)」
### 改动
**[web/core/step_impl/step7_standards.py](web/core/step_impl/step7_standards.py)** — `_IND_COMBINED` 加 `ind_401`
```python
"ind_401": {
"sub_ids": ("IND-401", "IND-402"),
"title": "IND-401 · 姓名校验",
"format": "字符集(汉字 + 少数民族 · + ASCII 字母)+ 长度 2-20",
},
```
**[web/core/orchestrator.py](web/core/orchestrator.py)** — 4 处改动
1. `_MERGED_SUBS` 加 `IND-401` / `IND-402`
2. `_HIDDEN_SUBS` 加 `IND-301` / `IND-302` / `IND-501` / `IND-502`
3. skip 分支的 `mapped_step` dict 加 `"IND-401": "std_ind_401"` / `"IND-402": "std_ind_401"`
4. `_COMBINED_STEP_META` 加 `std_ind_401`(order=301,身份信息组 300 起)
- `purpose` / `target` / `check` / `format` 都按新标题重写
**[web/configs/analysis_tree.json](web/configs/analysis_tree.json)** — `identity` 组
```json
{
"key": "identity",
"title": "身份信息",
"description": "IND-401 · 姓名校验(字符集 + 长度,2 子检查合一)",
"default_expand": true,
"children": [
{ "step_id": "std_ind_401" }
]
}
```
**[web/api/routes.py](web/api/routes.py)** — `_COMBINED_GROUPS` 加 `std_ind_401`
```python
_COMBINED_GROUPS = (
...
("std_ind_401", ("IND-401", "IND-402")),
)
```
→ 前端「字段匹配」输入框会聚合 `name / real_name / customer_name / xm / ...` + `姓名 / 客户名称 / ...` 注释关键字
### 烟测
```
=== Registered IND-001..016b / IND-4xx / IND-5xx / custom ===
std_ind_001 order=100 IND-001 · 身份证号校验
std_ind_002 order=101 IND-002 · 统一社会信用代码校验
std_ind_003 order=102 IND-003 · 手机号校验
std_ind_004 order=103 IND-004 · 行政区划代码校验
std_ind_005 order=104 IND-005 · 固定电话格式
std_ind_401 order=301 IND-401 · 姓名校验 ← 新合并
std_ind_016_b order=515 IND-016-b · 银行账号格式
(无 std_ind_301/302/501/502,无 std_ind_402)
=== /api/analysis-tree 模拟(hidden 过滤后)===
identity | ['std_ind_401'] ← 从 6 缩到 1
orphan: [] ← 无泄漏
=== match-config aggregation ===
std_ind_401 | names=['name','real_name','customer_name','user_name','xm',
'user_real_name','person_name']
| comments=['姓名','客户名称','用户名称','真实姓名','法人姓名',
'个人代理人姓名','经办人姓名','联系人姓名']
```
### 踩坑 / 注意
- **`sub_prefix` 推导对 `IND-401` / `IND-402` 这种没 `-a` 后缀的 ID 仍适用**:
`sub_ids[0].rsplit("-", 1)[0] + "-"` 对 `"IND-401"` → `"IND-"`,所以 `sub_short` 显示成 `401` / `402`(之前 a/b/c 是字母)。
error 字段前缀从 `[a] / [b]` 变成 `[401] / [402]`,用户视觉上是「多了个数字」但语义清楚。
- **`_derive_target` / `_derive_check` 里的 `IND-301 / -302 / -401 / -402 / -501 / -502` 特殊分支**:现在这几个 ID 都被 skip 了,那些分支变成 dead code,但留着不影响(恢复时还能用)。
- **没有匹配配置的损失**:IND-301 / -302 / -501 / -502 的 `applies_to_fields` / `comment_keywords` 在 YAML 里若有定义也暂时不被聚合到 `_COMBINED_GROUPS`,因为它们走的是隐藏路径而不是合并路径;后续要恢复只需从 `_HIDDEN_SUBS` 移除即可(orchestrator 自动重新注册为单 step)。
### 后续
- `_HIDDEN_SUBS` 集合已含 7 个 indicator(IND-016-a / -301 / -302 / -501 / -502 / -901 / -902)—— 命名规约继续稳定
- `_IND_COMBINED` 现含 5 个合并组(ind_001 / -002 / -003 / -004 / -401)—— 形态逐渐收敛
---
## 2026-08-13 · IND-016 / IND-901 / IND-902 隐藏 + 单位对公账户整组下架
### 需求
> 用户原始 2 轮反馈:
> 1. 「IND-016 的 a 和 b 检测的是同样的对象吗?」(确认 a 检测银行行号 CNAPS、b 检测银行账号,两个不同对象)
> 2. 「把 a 先隐藏,然后把 b 作为银行账号检测放到国标字段规范下面」
> 3. 「IND-901 和 IND-902 也隐藏,对公账户这整个大指标就不显示了」
### 设计决策
1. **IND-016-a(CNAPS 联行号)** 临时隐藏:暂不从 UI 跑
2. **IND-016-b(银行账号)** 移到「国标字段规范」组,与 IND-001 ~ IND-005 并列
3. **IND-901(单位名跨表唯一性)/ IND-902(账户名一致性)** 临时隐藏
4. **「单位对公账户 字段规范」整组** 删除(之前是 4 个子项:-a / -b / 901 / 902;现在 -b 走了,剩下 3 个都隐藏,组空了没必要留)
### 改动
**[web/core/orchestrator.py](web/core/orchestrator.py)** — 新增 `_HIDDEN_SUBS` 集合 + skip 逻辑
```python
# 2026-08-13 起临时隐藏(暂不从 UI 跑,按需恢复)
_HIDDEN_SUBS = (
"IND-016-a", # CNAPS 联行号:b 用,单位对公账户的大组已下架
"IND-901", # 单位名跨表唯一性
"IND-902", # 账户名一致性(跨字段)
)
# 跳过分支加:
if std_id in _HIDDEN_SUBS:
logger.info(f"跳过临时隐藏 indicator: {std_id}(UI 不展示,需要时从 _HIDDEN_SUBS 移除)")
continue
```
**[web/configs/analysis_tree.json](web/configs/analysis_tree.json)**:
- `standards` 组新增 `std_ind_016_b`,描述更新为「IND-001 ~ IND-016-b 系列:按 GB / GA / 央行 / ISO 标准做字段值级别合规校验(格式 / 校验位 / 出生日期 / 号段 / 编码存在性 / 银行账号 Luhn 等)」
- 删除整个 `business_unit_bank` 组
### 烟测
```
Total groups: 7
基础检查 children=3
国标字段规范 children=6 (新增 std_ind_016_b)
证件信息 children=5
身份信息 children=6
民政信息 children=3
业务字段规范 children=7
单位 / 法人 / 经办人 字段规范 children=7
std_ind_016_a: hidden
std_ind_016_b: in UI
std_ind_901: hidden
std_ind_902: hidden
```
### 后续
- `_HIDDEN_SUBS` 是临时方案:后续如果需要恢复,调一行就回来
- `IND-016-a` 的 indicator class 没动(保留 `standards/ind_016a_cnaps_format.py`),代码里 `load_standard_class("IND-016-a")` 仍能拉到
---
## 2026-08-13 · 国标字段规范描述精简 + 民政信息 / 证件信息 / 身份信息 默认展开
### 需求
1. 国标字段规范描述太长(约 161 字),展开后页面右侧出滚动条
2. 民政信息(以及 证件信息 / 身份信息)默认折叠,USER 抱怨「整个分析指标不显示」
### 改动
[web/configs/analysis_tree.json](web/configs/analysis_tree.json):
1. **`standards` 描述精简**:从 161 字 → 75 字
- 去掉「IND-001(身份证)/ IND-002(USCC)/ IND-003(手机号)/ IND-004(行政区划)多子检查已合并为单 step;IND-005(固定电话)单 sub」这段
- 合并后的下级指标点击看即可,描述里全列出来既冗余又占高度
2. **`id_info` / `identity` / `civil` 三组 `default_expand` 改为 `true`**
- 展开后页面会显示 8 个组,比 5 个组多 3 个,但都是单条小指标(3-6 个),整体高度可控
- 之前默认折叠让用户不知道有这些分析项
### 烟测
```
基础检查 expand=True desc_len=17 children=3
国标字段规范 expand=True desc_len=75 children=5
证件信息 expand=True desc_len=61 children=5
身份信息 expand=True desc_len=48 children=6
民政信息 expand=True desc_len=40 children=3
业务字段规范 expand=True desc_len=88 children=7
单位 / 法人 / 经办人 字段规范 expand=True desc_len=84 children=7
单位对公账户 字段规范 expand=True desc_len=66 children=4
```
8 个组全部默认展开,standards 描述 75 字(单行足够,不再出滚动条)。
---
## 2026-08-13 · IND-002 / -003 / -004 合并 + IND-005 改名(去 -a) ## 2026-08-13 · IND-002 / -003 / -004 合并 + IND-005 改名(去 -a)
### 需求 ### 需求
......
...@@ -272,6 +272,7 @@ async def get_match_config(): ...@@ -272,6 +272,7 @@ async def get_match_config():
("std_ind_002", ("IND-002-a", "IND-002-b", "IND-002-c")), ("std_ind_002", ("IND-002-a", "IND-002-b", "IND-002-c")),
("std_ind_003", ("IND-003-a", "IND-003-b")), ("std_ind_003", ("IND-003-a", "IND-003-b")),
("std_ind_004", ("IND-004-a", "IND-004-b")), ("std_ind_004", ("IND-004-a", "IND-004-b")),
("std_ind_401", ("IND-401", "IND-402")),
) )
for combined_step_id, sub_ids in _COMBINED_GROUPS: for combined_step_id, sub_ids in _COMBINED_GROUPS:
agg_names: list[str] = [] agg_names: list[str] = []
...@@ -304,6 +305,35 @@ async def get_match_config(): ...@@ -304,6 +305,35 @@ async def get_match_config():
"skip": False, "skip": False,
} }
# ── IND-006-b 不再带 -b 后缀:YAML key 仍叫 IND-006-b,UI step_id 改为 std_ind_006 ──
# 2026-08-13:IND-006-b(邮政编码)从「业务字段规范」移入「国标字段规范」组,
# step_id 与 IND-005 同样去掉后缀。
ind006_rec = step7.get("IND-006-b") or {}
ind006_atf = ind006_rec.get("applies_to_fields") or []
ind006_cmk = ind006_rec.get("comment_keywords") or []
by_step["std_ind_006"] = {
"names": ", ".join(ind006_atf),
"comments": ", ".join(ind006_cmk),
"skip": False,
}
# ── IND-013 的 3 个子检查改成新 IND 号:YAML key 仍叫 IND-013-a/b/c,UI step_id 改为 std_ind_017/018/019 ──
# 2026-08-13:原 IND-013-a/b/c(单位类型 / 经济类型 / 行业代码)从「单位/法人/经办人 字段规范」
# 移入「国标字段规范」组,step_id 各自分到 017/018/019(与 IND-005/006 同样去掉后缀)。
for src_id, target_step_id in (
("IND-013-a", "std_ind_017"),
("IND-013-b", "std_ind_018"),
("IND-013-c", "std_ind_019"),
):
rec = step7.get(src_id) or {}
atf = rec.get("applies_to_fields") or []
cmk = rec.get("comment_keywords") or []
by_step[target_step_id] = {
"names": ", ".join(atf),
"comments": ", ".join(cmk),
"skip": False,
}
# ── 其余 step(merge / empty / missing_comments 等)无需输入框 ── # ── 其余 step(merge / empty / missing_comments 等)无需输入框 ──
# 前端按 step_id 找不到时按 skip=True 处理即可 # 前端按 step_id 找不到时按 skip=True 处理即可
......
...@@ -15,94 +15,29 @@ ...@@ -15,94 +15,29 @@
{ {
"key": "standards", "key": "standards",
"title": "国标字段规范", "title": "国标字段规范",
"description": "IND-001 ~ IND-005 系列:按 GB / GA 标准做字段值级别合规校验。IND-001(身份证)/ IND-002(USCC)/ IND-003(手机号)/ IND-004(行政区划)多子检查已合并为单 step;IND-005(固定电话)单 sub。包含格式 / 校验位 / 出生日期 / 15 位兼容 / 号段 / 编码存在性 / 老代码兼容", "description": "IND-001 ~ IND-019 系列:按 GB / GA / 央行 / 国统字 / ISO 标准做字段值级别合规校验(格式 / 校验位 / 出生日期 / 号段 / 编码存在性 / 邮编 / 银行账号 Luhn / 单位类型 / 经济类型 / 行业代码 等)",
"default_expand": true, "default_expand": true,
"children": [ "children": [
{ "step_id": "std_ind_001" }, { "step_id": "std_ind_001" },
{ "step_id": "std_ind_002" }, { "step_id": "std_ind_002" },
{ "step_id": "std_ind_003" }, { "step_id": "std_ind_003" },
{ "step_id": "std_ind_004" }, { "step_id": "std_ind_004" },
{ "step_id": "std_ind_005" } { "step_id": "std_ind_005" },
] { "step_id": "std_ind_006" },
}, { "step_id": "std_ind_016_b" },
{ { "step_id": "std_ind_017" },
"key": "id_info", { "step_id": "std_ind_018" },
"title": "证件信息", { "step_id": "std_ind_019" }
"description": "IND-101 ~ IND-204:证件类型枚举 + 4 种证件号码(身份证 / 护照 / 港澳台居住证 / 外国人永居)",
"default_expand": false,
"children": [
{ "step_id": "std_ind_101" },
{ "step_id": "std_ind_201" },
{ "step_id": "std_ind_202" },
{ "step_id": "std_ind_203" },
{ "step_id": "std_ind_204" }
]
},
{
"key": "identity",
"title": "身份信息",
"description": "IND-301 ~ IND-502:证件类型+号码一致性 / 跨表唯一性 / 姓名 / 出生日期",
"default_expand": false,
"children": [
{ "step_id": "std_ind_301" },
{ "step_id": "std_ind_302" },
{ "step_id": "std_ind_401" },
{ "step_id": "std_ind_402" },
{ "step_id": "std_ind_501" },
{ "step_id": "std_ind_502" }
]
},
{
"key": "civil",
"title": "民政信息",
"description": "IND-601 ~ IND-603:行政区划代码 / 民族代码 / 婚姻状况代码",
"default_expand": false,
"children": [
{ "step_id": "std_ind_601" },
{ "step_id": "std_ind_602" },
{ "step_id": "std_ind_603" }
] ]
}, },
{ {
"key": "business", "key": "business_main",
"title": "业务字段规范", "title": "业务字段规范",
"description": "IND-006 ~ IND-011 系列:业务特定字段(通讯地址 / 邮编 / 个人公积金账号 / 首次参加工作年月 / 用工类型 / 本地户籍 / 职工状态);字段匹配走注释", "description": "IND-401 / IND-007-a:核心业务字段(姓名 / 个人公积金账号)",
"default_expand": true, "default_expand": true,
"children": [ "children": [
{ "step_id": "std_ind_006_a" }, { "step_id": "std_ind_401" },
{ "step_id": "std_ind_006_b" }, { "step_id": "std_ind_007_a" }
{ "step_id": "std_ind_007_a" },
{ "step_id": "std_ind_008_a" },
{ "step_id": "std_ind_009_a" },
{ "step_id": "std_ind_010_a" },
{ "step_id": "std_ind_011_a" }
]
},
{
"key": "business_unit",
"title": "单位 / 法人 / 经办人 字段规范",
"description": "IND-012 ~ IND-015 系列:单位基本信息(名称 / 类型 / 经济类型 / 行业代码 / 设立日期 / 启缴年月 / 发薪日);字段匹配走字段名 + 注释",
"default_expand": true,
"children": [
{ "step_id": "std_ind_012_a" },
{ "step_id": "std_ind_013_a" },
{ "step_id": "std_ind_013_b" },
{ "step_id": "std_ind_013_c" },
{ "step_id": "std_ind_014_a" },
{ "step_id": "std_ind_014_b" },
{ "step_id": "std_ind_015_a" }
]
},
{
"key": "business_unit_bank",
"title": "单位对公账户 字段规范",
"description": "IND-016 系列 + IND-901/902:开户银行行号 / 银行账号 格式 + 单位名跨表唯一性 + 账户名一致性(跨字段)",
"default_expand": true,
"children": [
{ "step_id": "std_ind_016_a" },
{ "step_id": "std_ind_016_b" },
{ "step_id": "std_ind_901" },
{ "step_id": "std_ind_902" }
] ]
} }
] ]
......
...@@ -551,6 +551,39 @@ def _register_indicator_steps() -> None: ...@@ -551,6 +551,39 @@ def _register_indicator_steps() -> None:
"IND-003-a", "IND-003-b", "IND-003-a", "IND-003-b",
"IND-004-a", "IND-004-b", "IND-004-a", "IND-004-b",
"IND-005-a", # 2026-08-13:去掉 -a 后缀,改名 std_ind_005 "IND-005-a", # 2026-08-13:去掉 -a 后缀,改名 std_ind_005
"IND-006-b", # 2026-08-13:去掉 -b 后缀,改名 std_ind_006,移入国标字段规范组
"IND-013-a", # 2026-08-13:去掉 -a 后缀,改名 std_ind_017,移入国标字段规范组
"IND-013-b", # 2026-08-13:去掉 -b 后缀,改名 std_ind_018
"IND-013-c", # 2026-08-13:去掉 -c 后缀,改名 std_ind_019
"IND-401", # 2026-08-13:与 IND-402 合并到 std_ind_401
"IND-402",
)
# 2026-08-13 起临时隐藏(暂不从 UI 跑,按需恢复)
_HIDDEN_SUBS = (
"IND-006-a", # 通讯地址:其他业务字段整组下架
"IND-008-a", # 首次参加工作年月:其他业务字段整组下架
"IND-009-a", # 用工类型:其他业务字段整组下架
"IND-010-a", # 本地户籍:其他业务字段整组下架
"IND-011-a", # 职工状态:其他业务字段整组下架
"IND-012-a", # 单位名称:单位/法人/经办人 字段规范整组下架
"IND-014-a", # 设立日期:单位/法人/经办人 字段规范整组下架
"IND-014-b", # 启缴年月:单位/法人/经办人 字段规范整组下架
"IND-015-a", # 发薪日:单位/法人/经办人 字段规范整组下架
"IND-016-a", # CNAPS 联行号:b 用,单位对公账户的大组已下架
"IND-101", # 证件类型枚举:证件信息整组下架
"IND-201", # 居民身份证 18 位:证件信息整组下架
"IND-202", # 中国护照:证件信息整组下架
"IND-203", # 港澳台居住证:证件信息整组下架
"IND-204", # 外国人永居证:证件信息整组下架
"IND-301", # 证件类型 + 证件号 一致性(跨字段):暂隐藏,按需恢复
"IND-302", # 证件唯一性(跨表):暂隐藏,按需恢复
"IND-501", # 出生日期格式:暂隐藏,按需恢复
"IND-502", # 出生日期年龄范围:暂隐藏,按需恢复
"IND-601", # 行政区划代码(民政部):暂隐藏,按需恢复
"IND-602", # 民族代码:暂隐藏,按需恢复
"IND-603", # 婚姻状况代码:暂隐藏,按需恢复
"IND-901", # 单位名跨表唯一性
"IND-902", # 账户名一致性(跨字段)
) )
prev_group: str | None = None prev_group: str | None = None
...@@ -572,11 +605,22 @@ def _register_indicator_steps() -> None: ...@@ -572,11 +605,22 @@ def _register_indicator_steps() -> None:
"IND-003-a": "std_ind_003", "IND-003-b": "std_ind_003", "IND-003-a": "std_ind_003", "IND-003-b": "std_ind_003",
"IND-004-a": "std_ind_004", "IND-004-b": "std_ind_004", "IND-004-a": "std_ind_004", "IND-004-b": "std_ind_004",
"IND-005-a": "std_ind_005", "IND-005-a": "std_ind_005",
"IND-006-b": "std_ind_006",
"IND-013-a": "std_ind_017",
"IND-013-b": "std_ind_018",
"IND-013-c": "std_ind_019",
"IND-401": "std_ind_401", "IND-402": "std_ind_401",
}.get(std_id, "?") }.get(std_id, "?")
logger.info( logger.info(
f"跳过子 indicator: {std_id}(已合并/改名到 {mapped_step} 复合 step)" f"跳过子 indicator: {std_id}(已合并/改名到 {mapped_step} 复合 step)"
) )
continue continue
# 跳过临时隐藏的子指标(2026-08-13 起,按需恢复)
if std_id in _HIDDEN_SUBS:
logger.info(
f"跳过临时隐藏 indicator: {std_id}(UI 不展示,需要时从 _HIDDEN_SUBS 移除)"
)
continue
group = meta.get("group", "通用") group = meta.get("group", "通用")
if group != prev_group: if group != prev_group:
same_group_count = 0 same_group_count = 0
...@@ -718,6 +762,27 @@ _COMBINED_STEP_META = { ...@@ -718,6 +762,27 @@ _COMBINED_STEP_META = {
), ),
"format": "GB/T 2260", "format": "GB/T 2260",
}, },
"std_ind_401": {
"parent_id": "ind_401",
"title": "IND-401 · 姓名校验",
"description": (
"姓名字段全维度校验:2 子检查合一(IND-401 字符集 + IND-402 长度)。"
"同一字段同一值多子检查失败的会合并成 1 行,error 字段拼全部子错误信息。"
),
"order": 301, # 身份信息组(300 起)
"purpose": (
"对姓名字段跑 IND-401(字符集:汉字 + 少数民族 · + ASCII 字母 + 空格 / 连字符 / 撇号,"
"禁止纯数字 / 纯特殊符号 / 控制字符 / Emoji)+ IND-402(trim 后长度 ∈ [2, 20])"
"2 个子检查,合并输出到 1 个 tab。"
),
"target": "字段名含 name / xm / username / 姓名 等姓名相关字段",
"check": (
"[401] 字符集:汉字 / 少数民族点 · / ASCII 字母 / 半角空格 / 连字符 / 撇号;"
"禁止纯数字 / 纯特殊符号 / 控制字符 / Emoji\n"
"[402] 长度:trim 后 ∈ [2, 20];纯符号 / 纯数字 / 长度 1 视为异常"
),
"format": "字符集 + 长度(无对应国标,业界通用经验值)",
},
} }
for _step_id, _meta in _COMBINED_STEP_META.items(): for _step_id, _meta in _COMBINED_STEP_META.items():
...@@ -782,6 +847,175 @@ register_step( ...@@ -782,6 +847,175 @@ register_step(
) )
# ── IND-006-b 改名(去掉 -b 后缀) + 移入国标字段规范组 ────────────
# 2026-08-13 起:原 IND-006-b(邮政编码 GB/T 23705-2009)从「业务字段规范」
# 移出 → 进「国标字段规范」,step_id 由 std_ind_006_b 改为 std_ind_006。
# IND-006-a(通讯地址)保留在「业务字段规范」不变(用户只动 b)。
# 与 IND-005-a → IND-005 改名同款模式:section_key / step_id / tab.key / tab.title 4 处同步改。
def _run_standards_ind_006(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides): # noqa: ARG001
from .step_impl.step7_standards import run_step7_for_indicator
res = run_step7_for_indicator(
"IND-006-b", cfg, dict_data=dict_data, log=log,
table_filter=table_filter,
match_override=match_overrides.get("std_ind_006") if match_overrides else None,
)
# 端到端重命名:section_key / step_id / tab.key / tab.title
res["section_key"] = "standards_ind_006"
proto = res.get("data", {}).get("_protocol") if isinstance(res.get("data"), dict) else None
if proto is not None:
proto["step_id"] = "std_ind_006"
for tab in (proto.get("tabs") or []):
if tab.get("key") == "standards_ind_006_b":
tab["key"] = "standards_ind_006"
if (tab.get("title") or "").startswith("IND-006-b"):
tab["title"] = "IND-006 · GB/T 23705 邮政编码格式"
return res
register_step(
step_id="std_ind_006",
title="IND-006 · 邮政编码格式",
description=(
"GB/T 23705 邮政编码格式校验(6 位数字 + 与行政区划代码前 2 位对应)。"
"原 std_ind_006_b 于 2026-08-13 去掉 -b 后缀、改名 std_ind_006,并从「业务字段规范」"
"组移入「国标字段规范」组。"
),
requires_db=True, llm_mode="none", required=False, order=105,
fn=_run_standards_ind_006,
detail=StepDetail(
purpose="对邮政编码字段跑格式校验:6 位数字;前 2 位对应省级行政区域(与 GB/T 2260 一致)。",
target="字段名含 zip / postal_code / postcode / youbian 或注释匹配的字段",
check=(
"a) 6 位纯数字(^\\d{6}$)\n"
"b) 前 2 位在 GB/T 2260 省级行政区划代码里(不强制县及县以上级)"
),
format="GB/T 23705-2009(格式)+ GB/T 2260(前 2 位映射)",
),
)
# ── IND-013-a / -b / -c 改名(去掉 -a/-b/-c 后缀)+ 移入国标字段规范组 ───
# 2026-08-13 起:原 IND-013 三个子检查各自分到一个全新的 IND 号:
# - IND-013-a(GB/T 12402 单位类别) → std_ind_017
# - IND-013-b(国统字〔2011〕86 号 经济类型) → std_ind_018
# - IND-013-c(GB/T 4754 行业代码) → std_ind_019
# 三个新号连续且不与现有 IND 号冲突(现有号:001-006/007-a..015-a,b/016-a,b/401 等)。
# 同 IND-005 / IND-006 改名同款模式:section_key / step_id / tab.key / tab.title 4 处同步改。
# 跑的还是原 IND-013-x 的 standard class(文件未动,registry 仍能找到)。
def _run_standards_ind_renamed(
*, source_id: str, target_step_id: str, target_title: str,
):
"""返回一个 runner:跑 source_id 的 indicator,端到端改名到 target_step_id / target_title。
闭包捕获 3 个参数:
- source_id: 原 indicator id(如 "IND-013-a")
- target_step_id: 目标 step_id(如 "std_ind_017")
- target_title: 目标 tab 标题(如 "IND-017 · GB/T 12402 单位类别 值域")
`log` 从 orchestrator dispatcher 传进来(不要在闭包里钉死,否则丢日志)
"""
target_section_key = f"standards_{target_step_id[len('std_'):]}" # "standards_ind_017"
def _runner(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides): # noqa: ARG001
from .step_impl.step7_standards import run_step7_for_indicator
res = run_step7_for_indicator(
source_id, cfg, dict_data=dict_data, log=log,
table_filter=table_filter,
match_override=match_overrides.get(target_step_id) if match_overrides else None,
)
# 端到端重命名:section_key / step_id / tab.key / tab.title
res["section_key"] = target_section_key
proto = res.get("data", {}).get("_protocol") if isinstance(res.get("data"), dict) else None
if proto is not None:
proto["step_id"] = target_step_id
for tab in (proto.get("tabs") or []):
# tab.key 原形:standards_ind_013_a → 改 target
if tab.get("key", "").startswith(f"standards_{source_id.lower().replace('-', '_')}"):
tab["key"] = target_section_key
# tab.title 原形:IND-013-a · GB/T 12402... → 改 target
if (tab.get("title") or "").startswith(source_id):
tab["title"] = target_title
return res
return _runner
# 3 个改名 step 复用同一 runner 工厂
register_step(
step_id="std_ind_017",
title="IND-017 · GB/T 12402 单位类别 值域",
description=(
"GB/T 12402-2017 单位类别分类与代码(2 位大类 / 4 位中类 / 文字形式)。"
"原 IND-013-a 于 2026-08-13 去掉 -a 后缀、改名 std_ind_017,并从「单位/法人/经办人 字段规范」"
"组移入「国标字段规范」组。"
),
requires_db=True, llm_mode="none", required=False, order=106,
fn=_run_standards_ind_renamed(
source_id="IND-013-a", target_step_id="std_ind_017",
target_title="IND-017 · GB/T 12402 单位类别 值域",
),
detail=StepDetail(
purpose="对单位类别字段跑值域校验:GB/T 12402-2017 接受 2 位大类码 / 4 位中类码 / 文字形式。",
target="字段名或注释含「单位类型 / 企业类型 / 单位性质 / 机构类型」",
check=(
"a) 接受 2 位大类码(10 内资 / 20 港澳台 / 30 外资 / 99 其他)\n"
"b) 接受 4 位中类码(公司 / 合伙 / 个体 / 其它细类)\n"
"c) 接受文字形式(公司 / 有限责任公司 / 国有企业 / 股份合作 等)"
),
format="GB/T 12402-2017",
),
)
register_step(
step_id="std_ind_018",
title="IND-018 · 国统字〔2011〕86 号 经济类型 值域",
description=(
"国统字〔2011〕86 号 经济类型分类与代码(3 位数字码 / 文字形式)。"
"原 IND-013-b 于 2026-08-13 去掉 -b 后缀、改名 std_ind_018,并从「单位/法人/经办人 字段规范」"
"组移入「国标字段规范」组。"
),
requires_db=True, llm_mode="none", required=False, order=107,
fn=_run_standards_ind_renamed(
source_id="IND-013-b", target_step_id="std_ind_018",
target_title="IND-018 · 国统字〔2011〕86 号 经济类型 值域",
),
detail=StepDetail(
purpose="对经济类型字段跑值域校验:国统字〔2011〕86 号 接受 3 位数字码 / 文字形式。",
target="字段名或注释含「经济类型 / 经济性质 / 登记注册类型 / 企业注册类型」",
check=(
"a) 接受 3 位数字码(100/110/120/.../390;含子项 + 兜底 190/290/390)\n"
"b) 接受文字形式(国有 / 集体 / 股份合作 / 联营 / 有限责任 / 股份 / 私营 / 中外合资 等)"
),
format="国统字〔2011〕86 号",
),
)
register_step(
step_id="std_ind_019",
title="IND-019 · GB/T 4754 行业代码 格式",
description=(
"GB/T 4754-2017 国民经济行业分类(门类 + 大类 / 中类 / 小类代码)。"
"原 IND-013-c 于 2026-08-13 去掉 -c 后缀、改名 std_ind_019,并从「单位/法人/经办人 字段规范」"
"组移入「国标字段规范」组。"
),
requires_db=True, llm_mode="none", required=False, order=108,
fn=_run_standards_ind_renamed(
source_id="IND-013-c", target_step_id="std_ind_019",
target_title="IND-019 · GB/T 4754 行业代码 格式",
),
detail=StepDetail(
purpose="对行业代码字段跑格式校验:GB/T 4754-2017 接受门类字母 + 2/3/4 位数字。",
target="字段名或注释含「行业代码 / 行业类别 / 行业分类 / 国民经济行业 / 所属行业」",
check=(
"a) 大类 3 位(门类字母 + 2 位数字,如 A01 农业)\n"
"b) 中类 4 位(门类字母 + 3 位数字,如 A011 谷物种植)\n"
"c) 小类 5 位(门类字母 + 4 位数字,如 A0111 稻谷种植)\n"
"d) 门类 ∈ A~T(20 个);数字段不以 00 开头\n"
"e) 仅校验格式;行业存在性留 IND-019-d(~2000 条表,加载不划算)"
),
format="GB/T 4754-2017",
),
)
register_step( register_step(
step_id="custom_value_check", step_id="custom_value_check",
title="自定义规则(字段值包含关键字)", title="自定义规则(字段值包含关键字)",
......
...@@ -1605,6 +1605,15 @@ _IND_COMBINED: dict[str, dict] = { ...@@ -1605,6 +1605,15 @@ _IND_COMBINED: dict[str, dict] = {
"title": "IND-004 · 行政区划代码校验", "title": "IND-004 · 行政区划代码校验",
"format": "GB/T 2260", "format": "GB/T 2260",
}, },
"ind_401": {
# 2026-08-13:IND-401 / IND-402 合并。注:这一对没有 -a/-b 后缀,
# 是直接以「姓名」「姓名长度」为 ID 的 2 个独立 indicator。
# sub_prefix 推导仍走得通("IND-401".rsplit("-", 1)[0] + "-" = "IND-"),
# 合并后 error 字段前缀显示 [401] / [402]。
"sub_ids": ("IND-401", "IND-402"),
"title": "IND-401 · 姓名校验",
"format": "字符集(汉字 + 少数民族 · + ASCII 字母)+ 长度 2-20",
},
} }
# 旧名保留以兼容 orchestrator 现有 import;新代码请用 _IND_COMBINED["ind_001"]["sub_ids"] # 旧名保留以兼容 orchestrator 现有 import;新代码请用 _IND_COMBINED["ind_001"]["sub_ids"]
......
Markdown is supported
0%
or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment