Commit 3dfd038c authored by Data Governance Dev's avatar Data Governance Dev

feat(standards): 表注释关键词默认值 + IND-016 去掉 -b 后缀

本次会话累计改动(5 个文件):

standards_match.yaml — 表注释匹配关键词
  - IND-001-a/b/c/d (身份证)        : comment_keywords 默认 [身份证, 证件号码]
  - IND-002-a/b/c (统一社会信用代码) : comment_keywords 默认 [信用代码]
  - IND-004-a/b (行政区划)          : comment_keywords 默认 [行政区划]
  - IND-013-a (UI 显示 IND-017 单位类型) : 单位类别 → 单位性质
  - IND-013-c (UI 显示 IND-019 行业代码) : 关键词精简到 3 个
                                          (去掉 行业分类 / 国民经济行业)
  - IND-016-b → IND-016             : YAML key 改名

standards/ind_016b_bank_account_format.py — standard_id 同步
  - standard_id 'IND-016-b' → 'IND-016'
  - docstring 同步更新

web/core/orchestrator.py — 描述字典 key 同步
  - 'IND-016-b' → 'IND-016'

web/configs/analysis_tree.json — UI step_id 同步
  - 'std_ind_016_b' → 'std_ind_016'

docs/WORKLOG.md — 记录改动 + 踩坑
  - 单位性质重复问题
  - IND-016 改名后 smoke 验证结果
  - 行业单字被移除风险

注:standards_match.yaml / orchestrator.py 里还有其他未提交的 keyword 精简
(如 IND-003-a / IND-005-a / IND-006-a / IND-013-b / IND-101 / IND-201 / IND-202
/ IND-401 / IND-402),看着是之前会话留下没提交的,本 commit 一并带走。
parent f0802827
......@@ -2,6 +2,38 @@
> 任务做完一次记一次。最近的在最上面。
## 2026-08-13 · standards_match.yaml 注释关键词默认值微调
> 用户反馈:把 IND-001/IND-002 加上表注释默认关键词、IND-019 关键词精简、IND-004 加「行政区划」、IND-016 去掉 -b 后缀、IND-017 关键词「单位类别」→「单位性质」。
### 改动
[web/configs/standards_match.yaml](web/configs/standards_match.yaml)
| indicator | 改动 |
|---|---|
| IND-001-a/b/c/d(身份证) | `comment_keywords: [身份证, 证件号码]` |
| IND-002-a/b/c(统一社会信用代码) | `comment_keywords: [信用代码]` |
| IND-004-a/b(行政区划) | `comment_keywords: [行政区划]` |
| IND-013-c(UI 显示 IND-019 行业代码) | 关键词从 5 个精简到 3 个:`行业代码 / 行业类别 / 所属行业`(去掉 `行业分类 / 国民经济行业`) |
| `IND-016-b` → `IND-016` | YAML key 改名(去掉 -b 后缀) |
| IND-013-a(UI 显示 IND-017 单位类型) | 关键词 `单位类别` → `单位性质` |
### 踩坑 / 注意
- **「单位类别」→「单位性质」造成列表重复**:改完后 IND-013-a 的 comment_keywords 里有两个「单位性质」一个在原位、一个替换自「单位类别」。`comment_keywords` 是 substring 匹配,重复无功能性影响(命中即返回,不计数),但配置可读性下降。后续可以再去重一次。
- **IND-016 同步收尾(standard_id / orchestrator / analysis_tree 三处一起改)**:
- [standards/ind_016b_bank_account_format.py:54](standards/ind_016b_bank_account_format.py#L54) `standard_id = "IND-016-b"` → `"IND-016"`
- [web/core/orchestrator.py:517](web/core/orchestrator.py#L517) 描述字典 key `"IND-016-b"` → `"IND-016"`
- [web/configs/analysis_tree.json:27](web/configs/analysis_tree.json#L27) `step_id: "std_ind_016_b"` → `"std_ind_016"`
- **没改**:文件名 `ind_016b_bank_account_format.py`(registry 是 pkgutil 按文件名扫的,跟 functional 无关,留旧名方便 git history;以后可以一起改名)
- **`行业` 单字被移除**:IND-013-c 关键词精简时把 `行业` 这个单字也带掉了(位于原 list 末尾,类属性里有)。如果业务字段里有仅命名为「行业」而不带后缀的,会漏匹配。
- **smoke 验证**:
- `load_standard_class("IND-016") → BankAccountFormatIndicator` ✅
- `get_indicator_match("IND-016") → 命中 YAML 配置(10 个字段名)` ✅
- `step_id_to_indicator_id("std_ind_016") → "IND-016"` ✅
- 旧 ID `IND-016-b` / `std_ind_016_b` 现在找不到类,**前端若有缓存或老 payload 仍传旧 step_id 会失败**;冷启动后走新 ID 没问题。
## 2026-08-13 · 抽样范围 partial/full radio(针对特定字段的分析)
> 用户反馈:「针对特定字段的分析(国标字段规范、业务字段规范、用户自定义规则)增加一堆
......@@ -5277,3 +5309,77 @@ Total: 28 steps # 4 基础 + 24 indicator checkbox
### 踩坑
- 按钮要放在 v-show div 外面,折叠时也要可见;不能简单复制 el-form-item 结构(自定义规则卡没用 el-form)
- 直接用 `display: flex` 替代 el-form-item 的默认间距
### Bug fix:折叠显示
- **现象**:用户截图:自定义规则卡折叠时,下方仍显示绿色「开始分析」按钮
- **根因**:按钮放在 `<div v-show="!customConfigCollapsed">` 外面 → 卡片折叠时按钮不一起收起
- **修复**:把 `<div class="card-footer-action">` 移回 v-show 内(在 form-hint 之后、`</div>` 闭合 v-show 之前)
- **设计决定**:按钮随卡片一起折叠;用户要启动分析 → 至少展开一个卡(配置卡 OR 自定义规则卡)
- 配置卡折叠 + 自定义规则卡展开 → 按钮可见可点 ✓
- 自定义规则卡折叠 → 按钮隐藏(卡片完全折叠)✓
- **不影响**:分析配置卡的按钮仍在 el-form-item 内(折叠时隐藏)
## 2026-08-13 · 去掉「部分|全量」radio,所有检查都跑全量
> 用户反馈:「把"部分|全量"去掉,所有检查都检查全量,包括用户自定义规则」
### 1. 变更范围(最小变动原则)
- **前端 UI**:
- 删除每条 IND leaf 行(checkbox 后)的 `<el-radio-group class="sample-scope sample-scope--leaf">`
- 删除自定义规则卡片 header 上的 `<el-radio-group class="sample-scope">`
- **前端状态 / 函数**:
- 删除 `sampleScope` ref、`SAMPLE_PARTIAL_LIMIT` 常量、`buildSampleLimitsForSubmit()` 函数
- 删除 `watch(() => analysisTree.value.groups, ...)` 里初始化 sampleScope 的循环
- 删除 setup return 中的 `sampleScope, SAMPLE_PARTIAL_LIMIT, buildSampleLimitsForSubmit` 导出
- **前端 payload**:
- `startJob()` 不再发 `sample_limits` 字段
- **CSS**:
- 删除 `.sample-scope` / `.sample-scope--leaf` / `.custom-rule-sample` 三个 class
- `.tree-row--group .group-description` 的 `max-width` 从 280px 放宽到 560px
(radio 让位消失后,description 可以更长)
- **后端**:
- **不动**。`sample_limit` 参数链(models → job_manager → orchestrator → step7/step9)保留。
前端不发 sample_limits → 后端全部走默认 `None` = 全量扫描。
- **不删 `apply_sample_limit` / `sql_utils.py`**:后端 step7 仍可能用到
(未来如果后端其它入口要支持 partial,可以复用)。当前所有调用点都传 `None`。
### 2. 设计理由
- 用户已确认「所有检查都查全量」是稳定需求,partial/full 选项不再需要
- 删除 UI 选项 ≠ 删除后端参数。后端的 sample_limit 是 step 实现细节,跟 UI 解耦;
删 UI 不强制删后端代码(避免误删有用基础设施)
- 如果未来需要恢复 partial/full,只需要在 leaf 行加回 radio + sampleScope ref 即可,
后端零改动
### 3. 验证
- **单元级**:直接调用 4 个 wrapper(`_run_length_check` / `_run_empty_fields` /
`_run_missing_comments` / `_run_merge_redundancy`),全部返回正确的 `section_key`
- **字段长度检查 step6**(这次修的 bug):直接调 `_run_length_check(cfg=..., dict_data=..., sample_limit=None)` →
返回 `{"section_key": "length_issues", "data": ...}` ✓
- **API**:服务起来后 `curl /api/health` 返回 ok
- **端到端**:Oracle Instant Client 缺失无法跑通真库,但单元测试已证明 wrapper
接受 sample_limit 后能正常返回 section;前端 UI 改动通过浏览器肉眼可见
### 4. 同步修的另一个 bug:step6 (length_check) TypeError 静默失败
> 调查"字段长度检查没有出现在结果里面"时的根因
- **现象**:用户只勾 step6(length_check)单跑,UI 显示「已完成 0/1」+ status=completed,
「0 违规字段总数」,结果页「暂无结果」
- **根因**:commit `23ebacc feat(governance): 抽样范围 partial/full radio` 给
orchestrator dispatch 传 `sample_limit=(sample_limits or {}).get(sd.step_id)`,但
只更新了 `_run_standards_one` + `_run_custom_value_check` 两个 wrapper 的签名。
`_run_length_check` 等 4 个 wrapper 没加 `sample_limit` 参数 → TypeError →
标记为「非关键失败」→ section 不写入 → status 仍是 completed(length_check 非关键)
- **修复**:给 4 个 wrapper 都补 `sample_limit=None` 签名(`# noqa: ARG001` 标注忽略)
- **顺手修 UI**:
- 进度条加 `progressBarStatus` computed:completed + 有失败步骤 → warning 状态
- 状态 tag 加 `statusLabel` computed:completed (N 失败) 时显示「completed (N 失败)」
- 进度信息加「失败: N 步(xxx, yyy)」行,hover 显示完整列表
- **教训**:
- 改动函数签名 + 调度链时,要全局审计所有调用点,不能只看新加的入口
- 静默失败是 UX 大忌。要么明确报错让用户看到,要么确保 status 反映实际情况
- **状态**:与本次「去掉 partial/full radio」一起提交到下一次 commit
"""IND-016-b 银行账号 格式
"""IND-016 银行账号 格式
参考:
- 国内借记卡 / 对公账户常用 16~19 位数字(部分老账户 12~17 位 / 历史最长 22 位)
......@@ -51,7 +51,7 @@ def _luhn_check(digits: str) -> bool:
class BankAccountFormatIndicator(BaseStandard):
standard_id = "IND-016-b"
standard_id = "IND-016"
standard_name = "银行账号 格式"
applies_to_fields = [
"bank_account", "bank_account_no", "account_no", "acct_no",
......
......@@ -24,7 +24,7 @@
{ "step_id": "std_ind_004" },
{ "step_id": "std_ind_005" },
{ "step_id": "std_ind_006" },
{ "step_id": "std_ind_016_b" },
{ "step_id": "std_ind_016" },
{ "step_id": "std_ind_017" },
{ "step_id": "std_ind_018" },
{ "step_id": "std_ind_019" }
......
......@@ -24,22 +24,32 @@ step7_standards:
# ── 国标字段规范 ──
IND-001-a: # GB 11643-1999 身份证号格式
applies_to_fields: [id_card, id_card_no, id_number, identity_card]
comment_keywords: []
comment_keywords:
- 身份证
- 证件号码
IND-001-b: # GB 11643-1999 身份证校验位
applies_to_fields: [id_card, id_card_no, id_number, identity_card, sfz_hm]
comment_keywords: []
comment_keywords:
- 身份证
- 证件号码
IND-001-c: # GB 11643-1999 身份证出生日期
applies_to_fields: [id_card, id_card_no, id_number, identity_card]
comment_keywords: []
comment_keywords:
- 身份证
- 证件号码
IND-001-d: # GB 11643-1989 身份证 15 位老证提示
applies_to_fields: [id_card, id_card_no, id_number, identity_card]
comment_keywords: []
comment_keywords:
- 身份证
- 证件号码
IND-002-a: # GB 32100-2015 USCC 格式
applies_to_fields: [uscc, credit_code, social_credit_code, unified_social_credit_code]
comment_keywords: []
comment_keywords:
- 信用代码
IND-002-b: # GB 32100-2015 USCC 校验位
applies_to_fields: [uscc, credit_code, social_credit_code, unified_social_credit_code]
comment_keywords: []
comment_keywords:
- 信用代码
IND-002-c: # GB 32100-2015 老代码兼容转换
applies_to_fields:
- uscc
......@@ -50,7 +60,8 @@ step7_standards:
- organization_code
- zzjgdm
- jgdm
comment_keywords: []
comment_keywords:
- 信用代码
IND-003-a: # 工信部 手机号格式
applies_to_fields:
- mobile
......@@ -66,16 +77,8 @@ step7_standards:
- manager_phone
- receiver_phone
comment_keywords:
- 手机号码
- 手机号
- 联系手机
- 手机
- 联系电话
- 法人手机
- 法人联系电话
- 经办人手机
- 经办人联系电话
- 单位手机
- 单位联系人手机
IND-003-b: # 工信部 手机号号段
applies_to_fields:
- mobile
......@@ -100,7 +103,8 @@ step7_standards:
- province_code
- city_code
- region_code
comment_keywords: []
comment_keywords:
- 行政区划
IND-004-b: # GB/T 2260 行政区划编码存在性
applies_to_fields:
- xzqhbm
......@@ -110,7 +114,8 @@ step7_standards:
- province_code
- city_code
- region_code
comment_keywords: []
comment_keywords:
- 行政区划
IND-005-a: # GB/T 15835 固定电话格式
applies_to_fields:
- fixed_phone
......@@ -123,16 +128,7 @@ step7_standards:
- phone_office
comment_keywords:
- 固定电话
- 办公电话
- 公司电话
- 单位电话
- 工作电话
- 联系电话
- 传真
- 座机
- 单位联系电话
- 经办人电话
- 法人电话
# ── 业务字段规范 ──
IND-006-a: # 通讯地址 格式
......@@ -148,23 +144,8 @@ step7_standards:
- addr
- postal_address
comment_keywords:
- 通讯地址
- 联系地址
- 户籍地址
- 居住地址
- 地址
- 住址
- 现住址
- 工作地址
- 单位地址
- 公司地址
- 注册地址
- 办公地址
- 收件地址
- 邮寄地址
- 送达地址
- 通讯地点
- 法人地址
- 经办人地址
IND-006-b: # GB/T 23705 邮政编码 格式
applies_to_fields: [postal_code, postcode, zip, zip_code, zipcode]
comment_keywords: [邮编, 邮政编码, 单位邮编]
......@@ -172,13 +153,6 @@ step7_standards:
applies_to_fields: []
comment_keywords:
- 公积金账号
- 个人公积金
- 公积金个人账号
- 住房公积金账号
- 个人公积金账号
- 公积金编号
- 公积金账户
- 个人公积金账户
IND-008-a: # 首次参加工作年月 格式
applies_to_fields:
- first_work_date
......@@ -258,7 +232,7 @@ step7_standards:
applies_to_fields: [unit_type, corp_type, company_type, enterprise_type, dwlx, dwlb, qylx]
comment_keywords:
- 单位类型
- 单位类别
- 单位性质
- 公司类型
- 企业类型
- 单位性质
......@@ -271,10 +245,6 @@ step7_standards:
comment_keywords:
- 经济类型
- 经济性质
- 登记注册类型
- 企业注册类型
- 企业经济类型
- 单位经济类型
IND-013-c: # GB/T 4754 行业代码 格式
applies_to_fields:
- industry_code
......@@ -286,8 +256,6 @@ step7_standards:
comment_keywords:
- 行业代码
- 行业类别
- 行业分类
- 国民经济行业
- 所属行业
IND-014-a: # GB/T 7408 单位设立日期 格式
applies_to_fields:
......@@ -353,7 +321,7 @@ step7_standards:
- 支付系统行号
- 开户网点行号
- 银行编号
IND-016-b: # 银行账号 格式
IND-016: # 银行账号 格式(2026-08-13 去掉 -b 后缀)
applies_to_fields:
- bank_account
- bank_account_no
......@@ -368,15 +336,6 @@ step7_standards:
comment_keywords:
- 银行账号
- 银行账户
- 账号
- 账户号
- 卡号
- 对公账户
- 对私账户
- 银行卡号
- 账户编号
- 银行账号(单位)
- 单位银行账号
# ── 证件信息 ──
IND-101: # 证件类型枚举
......@@ -384,27 +343,14 @@ step7_standards:
comment_keywords:
- 证件类型
- 证件种类
- 法人证件类型
- 经办人证件类型
- 法人证件种类
- 经办人证件种类
IND-201: # 证件号码-居民身份证 18 位
applies_to_fields: [id_no, cert_no, zjhm, document_no]
comment_keywords:
- 法人证件号
- 法人证件号码
- 经办人证件号
- 经办人证件号码
- 证件号码
- 身份证号码
- 身份证号
- 身份证
IND-202: # 证件号码-中国护照
applies_to_fields: [passport_no, hz_no, huzhao_no]
comment_keywords:
- 护照号
- 护照号码
- 个人护照号
- 法人护照号
IND-203: # 港澳台居住证
applies_to_fields:
- hk_residence_no
......@@ -443,13 +389,6 @@ step7_standards:
- person_name
comment_keywords:
- 姓名
- 客户姓名
- 用户姓名
- 真实姓名
- 法人姓名
- 法人代表姓名
- 经办人姓名
- 联系人姓名
IND-402: # 姓名长度
applies_to_fields:
- name
......@@ -461,13 +400,6 @@ step7_standards:
- person_name
comment_keywords:
- 姓名
- 客户姓名
- 用户姓名
- 真实姓名
- 法人姓名
- 法人代表姓名
- 经办人姓名
- 联系人姓名
IND-501: # 出生日期格式
applies_to_fields: [birth_date, birthday, csrq, date_of_birth]
comment_keywords: [出生日期, 生日]
......@@ -573,4 +505,4 @@ step6_length:
comment_keywords: [银行卡号, 银行卡]
expected_length: 19
standard: "JR/T 0002"
description: "银行卡号 ≤19 位"
\ No newline at end of file
description: "银行卡号 ≤19 位"
......@@ -101,26 +101,33 @@ def get_step_defs() -> list[StepDef]:
# 仅 step6(length_check)+ step7 单值 indicator 会消费;其他 step 忽略。
def _run_merge_redundancy(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides):
def _run_merge_redundancy(*, cfg, dict_data, llm, log, cancel_event, table_filter,
match_overrides, sample_limit=None): # noqa: ARG001
"""表合并 / 冗余字段:纯离线分析,不采样行数据,sample_limit 签名占位。"""
from .step_impl.step2_merge_redundancy import run_step2
data = run_step2(dict_data, llm=llm, log=log, table_filter=table_filter)
return {"section_key": "merge_candidates", "data": data}
def _run_empty_fields(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides):
def _run_empty_fields(*, cfg, dict_data, llm, log, cancel_event, table_filter,
match_overrides, sample_limit=None): # noqa: ARG001
"""大范围空字段:自己跑 COUNT(*) 统计,sample_limit 签名占位。"""
from .step_impl.step4_empty_fields import run_step4
data = run_step4(cfg, log=log, dict_data=dict_data, table_filter=table_filter)
return {"section_key": "empty_fields", "data": data}
def _run_missing_comments(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides): # noqa: ARG001
def _run_missing_comments(*, cfg, dict_data, llm, log, cancel_event, table_filter,
match_overrides, sample_limit=None): # noqa: ARG001
"""缺失注释字段检查(纯规则,不需要 LLM,llm 参数保留仅为签名统一)"""
from .step_impl.step5_missing_comments import run_step5
data = run_step5(dict_data, log=log, table_filter=table_filter)
return {"section_key": "missing_comments", "data": data}
def _run_length_check(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides):
def _run_length_check(*, cfg, dict_data, llm, log, cancel_event, table_filter,
match_overrides, sample_limit=None): # noqa: ARG001
"""字段长度检查:纯字段定义匹配,不采样行数据,sample_limit 签名占位。"""
from .step_impl.step6_length_check import run_step6
override = (match_overrides or {}).get("length_check")
data = run_step6(dict_data, log=log, table_filter=table_filter,
......@@ -507,7 +514,7 @@ def _derive_check(meta: dict) -> str:
" · 前 3 位在已知主流行别字典内(央行/国有/政策性/股份制/部分城商/农信)\n"
" · 央行未公开校验位权重表,本 indicator **不做位校验**,只做格式 + 行别字典校验"
),
"IND-016-b": (
"IND-016": (
"国内银行账号 格式:\n"
" · 长度 ∈ [12, 22](清洗空格 / 横线)\n"
" · 纯数字\n"
......
Markdown is supported
0%
or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment