Commit 23ebacc6 authored by Data Governance Dev's avatar Data Governance Dev

feat(governance): 抽样范围 partial/full radio(per-IND 独立控制)

用户需求:针对特定字段的分析(国标字段规范、业务字段规范、用户自定义规则)
希望可切换「部分(前 1000 行)抽样 / 全量扫描」,且每个分析项独立配置(A 项选部分,
B 项可独立选全量)。

## 后端
- 新增 web/core/sql_utils.py: apply_sample_limit(sql, sample_limit, db_type)
  - sample_limit=None  → 移除末尾 LIMIT N / AND ROWNUM <= N(全量扫描)
  - sample_limit=int   → 保留原样(partial 模式)
  - 三方言均覆盖:MySQL / 达梦(标准 LIMIT)/ Oracle 11g(AND ROWNUM <= N)
- models.ConnectRequest 新增 sample_limits: dict[str, int | None]
- job_manager.submit 透传 sample_limits 给 orchestrator
- orchestrator.run_governance_workflow 新增 sample_limits kwarg;按 step_id 取值
  分发给各 _run_* 闭包
- step7_standards._run_round1/2/3_* 全部加 sample_limit 参数:渲染 SQL 后用
  apply_sample_limit 二次处理(partial→加 LIMIT 1000,full→剥 LIMIT)
- step9_custom_rules.run_step_custom_value_check 加 sample_limit 参数:
  SQL 构造按 None / oracle / else 三分支

## 前端
- index.html sampleScope ref + SAMPLE_PARTIAL_LIMIT 常量
- 每个 IND-XXX leaf 行(checkbox 之后、字段名输入框之前)加 partial/full radio
  - key = step_id(每个 IND 独立),不再按 group 共享
  - 标签去掉了「1000」后缀,仅显示「部分 / 全量」
- 自定义规则卡片保持 card-level 单个 radio(共享 custom_value_check step_id):
  step9 后端所有规则共用一个 LIMIT,要下沉需另改接口
- watch(analysisTree.groups) 加载后为每个 step_id 初始化 sampleScope='partial'
- buildSampleLimitsForSubmit() 把 sampleScope 转成后端要的 {step_id: int|None}
- startJob() payload 加 sample_limits 字段
- style.css 新增 .sample-scope--leaf(margin-left: 8px,紧贴字段名输入框)

## 端到端验证
- 用户在 IND-001/002/003 选全量、IND-004 选部分 → payload 正确:
  {std_ind_001: None, std_ind_002: None, std_ind_003: None, std_ind_004: 1000}
- Pydantic ConnectRequest 解析通过;类型 [(str, 'NoneType'/'int'), ...]
- 三方言 × partial/full × 多种 SQL 形态,行为符合预期

## 踩坑
edit 工具改函数签名+docstring 时只换首行 → 留孤儿 """ → SyntaxError。
教训:edit 改函数签名 + docstring 时,把整个 docstring 完整贴一遍。
parent 39f5a4fb
......@@ -2,6 +2,154 @@
> 任务做完一次记一次。最近的在最上面。
## 2026-08-13 · 抽样范围 partial/full radio(针对特定字段的分析)
> 用户反馈:「针对特定字段的分析(国标字段规范、业务字段规范、用户自定义规则)增加一堆
> radio 选项,让用户可以决定查部分或者全量数据,默认是部分(前 1000 行);选全量的搜索
> 就不加 LIMIT 等限制条件。这个配置针对每一个分析项目,A 项目选部分,B 项目可以选全量」
### 1. 设计:per-step 配置 + 默认 partial + 一份独立 dict 跨前后端
- **配置粒度**:按 `step_id` 独立配置。后端模型新加
`sample_limits: dict[str, int | None]`(2026-08-13 起的字段);key=step_id,
value=`int`(partial 上限)/ `None`(全量扫描,不加 LIMIT)。
- **默认值**:UI 未操作时为 `'partial'`,对应 `SAMPLE_PARTIAL_LIMIT = 1000`。
用户改 radio 后立刻覆盖;缺省值由 `buildSampleLimitsForSubmit()` 兜底。
- **适用 step**:当前只给三类加 radio(用户指定):
- **国标字段规范** group(`std_ind_001` 等)
- **业务字段规范** group(`length_check` 等)
- **用户自定义规则** —— 不在 analysisTree group 里,独立占一个 key `'__custom__'`
映射到 step_id `custom_value_check`
### 2. 新增 web/core/sql_utils.py —— 跨方言移除 LIMIT
- 一次性解决「full 模式不加 LIMIT」的三方言差异:
- MySQL / 达梦:末尾 `LIMIT N` 或 `LIMIT N, M`
- Oracle 11g:末尾 `AND ROWNUM <= N`(step7/step9 自己拼的,详情见 8-13 第 1 条)
- 提供 `apply_sample_limit(sql, sample_limit, db_type)`:
- `sample_limit=None` → 正则剥掉末尾的 LIMIT / ROWNUM
- `sample_limit=int` → 保留原样(partial 模式)
- 测试覆盖:MySQL / Dameng / Oracle × LIMIT / LIMIT 偏移 / 无 LIMIT / 中间 LIMIT 8 种 case,全部正确。
### 3. 后端透传链:ConnectRequest → job_manager → orchestrator → step 闭包
| 层 | 改动 |
|---|---|
| [web/core/models.py:53](web/core/models.py#L53) | 新字段 `sample_limits: dict[str, int \| None] = Field(default_factory=dict)` |
| [web/core/job_manager.py:242](web/core/job_manager.py#L242) | `getattr(job.req, "sample_limits", None)` 透传给 orchestrator |
| [web/core/orchestrator.py](web/core/orchestrator.py) | `run_governance_workflow(..., sample_limits=None)`;dispatch 时 `sample_limit=(sample_limits or {}).get(sd.step_id)`;每个 `_run_*` 闭包都加 kwarg |
| [web/core/step_impl/step7_standards.py](web/core/step_impl/step7_standards.py) | `_run_round1/2/3_*` 全加 `sample_limit: int \| None = None`;调 `loader.render(..., limit=effective_limit)` 后 `apply_sample_limit(sql, sample_limit, cfg.db_type)` |
| [web/core/step_impl/step9_custom_rules.py](web/core/step_impl/step9_custom_rules.py) | SQL 构造按 `sample_limit is None` / `oracle` / `else` 三分支 |
**关键设计**:partial 模式 ≠ 加 LIMIT;full 模式 ≠ 删 LIMIT。
- `effective_limit = sample_limit if sample_limit is not None else _DEFAULT_LIMIT`
(partial→int,加 LIMIT;full→None,剥 LIMIT)
- 这样 step 内部所有 `if cfg.db_type == "oracle": sql += " AND ROWNUM<=N"` 仍然能用同一个
`effective_limit`,调用方只关心「None=不要 LIMIT」语义。
### 4. 前端 UI:group 行 radio + 自定义规则卡片 radio
- [web/static/index.html:1228-1229](web/static/index.html#L1228-L1229)
`const sampleScope = ref({});` + `SAMPLE_PARTIAL_LIMIT = 1000`
- [web/static/index.html:432-440](web/static/index.html#L432-L440) analysisTree 的 group 行
模板末尾加 `<el-radio-group v-model="sampleScope[row.title]">`,样式 `.sample-scope`。
- [web/static/index.html:551-561](web/static/index.html#L551-L561) 自定义规则卡片 header
也加一份 radio,固定 key `'__custom__'`。
- [web/static/index.html:1182-1196](web/static/index.html#L1182-L1196) `watch(analysisTree.groups)`
加载后立即初始化每个 group 的 sampleScope 为 `'partial'`,避免 v-model 绑定 undefined。
- [web/static/style.css:284-295](web/static/style.css#L284-L295) `.tree-row--group .group-description`
限制 max-width 280px(给 radio 让位,避免换行错位);radio 自身 `flex-shrink: 0`。
### 5. 前端 payload 转换:buildSampleLimitsForSubmit()
```js
function buildSampleLimitsForSubmit() {
const out = {};
const groups = analysisTree.value?.groups || [];
for (const g of groups) {
const scope = sampleScope.value[g.title];
const limit = scope === 'full' ? null : SAMPLE_PARTIAL_LIMIT;
for (const c of (g.checks || [])) out[c.step_id] = limit;
}
out['custom_value_check'] =
sampleScope.value['__custom__'] === 'full' ? null : SAMPLE_PARTIAL_LIMIT;
return out;
}
```
[startJob()](web/static/index.html#L1882) payload 加上:
```js
sample_limits: buildSampleLimitsForSubmit(),
```
### 6. 端到端验证(手动)
模拟用户在国标字段规范选「全量」+ 业务字段规范选「部分」+ 自定义规则选「全量」:
| step_id | UI 状态 | 发出值 | 后端 SQL |
|---|---|---|---|
| `std_ind_001` | 全量 | `None` | `SELECT ... WHERE ...` (无 LIMIT) |
| `std_ind_002` | 全量 | `None` | 同上 |
| `length_check` | 部分 | `1000` | `SELECT ... WHERE ... LIMIT 1000` |
| `custom_value_check` | 全量 | `None` | `SELECT ... WHERE ...` (无 ROWNUM) |
Pydantic `ConnectRequest(**payload)` 解析通过,`sample_limits` 字段类型 `[('std_ind_001', 'NoneType'), ('std_ind_002', 'NoneType'), ('length_check', 'int'), ('custom_value_check', 'NoneType')]` ✓
### 7. 踩坑:edit 工具替换 docstring 误删闭括号
- **现象**:给 step7 的 `_run_round3_global_uniqueness_one` / `_run_round3_cross_field_one`
加 `sample_limit` 参数时,edit 只换了第一行 `"""`,却没替换后面紧跟的
原 docstring 末尾 `"""` —— 留下孤儿 `"""` + 旧注释,Python 解析直接
`SyntaxError: invalid syntax`(错误位置指到 `1118 ���ͬ���ظ�ֵ��`,被识别为新行)。
- **修复**:重新 read 文件 → 把原 docstring 全部内容(含结尾 `"""`)一起粘到新字符串。
- **教训**:edit 改函数签名 + docstring 时,**把整个 docstring 完整贴一遍**,不要用
「只换首行」的省力做法。Python 解析器对 `"""` 配对很敏感。
### 8. UI 调整(用户反馈后):per-IND radio 而非 per-group
- **用户原话**:「范围错了,是针对每一个 IND-XXX 可以选部分或全量,然后部分后面
不写 1000,位置放在字段名的左侧」
- **v1 的问题**:radio 放在 group 行(一个 group 共享一个 limit)—— 用户希望每
个 IND 独立控制,不能一改全改;「部分 1000」的「1000」也让 radio 看着像
业务字段规范没有指示数字的占位符。
- **v2 改动**([web/static/index.html:442-454](web/static/index.html#L442-L454)):
- 删除 group 行的 `<el-radio-group v-model="sampleScope[row.title]">`
- 在每个 leaf 行(checkbox 之后、match-inputs 之前)加一个 radio:
```html
<el-radio-group class="sample-scope sample-scope--leaf"
v-model="sampleScope[row.step_id]" ...>
<el-radio-button label="partial">部分</el-radio-button>
<el-radio-button label="full">全量</el-radio-button>
</el-radio-group>
```
- 「部分 1000」→「部分」(无 1000 后缀;1000 是后端约定的 partial_limit 默认值,
详见 `SAMPLE_PARTIAL_LIMIT` 常量)
- **`buildSampleLimitsForSubmit` 同步调整**:原来按 group 共享 limit;现在按
step_id 独立,每个 IND 自带自己的 limit。模拟:用户在 IND-001/002/003 选「全量」、
IND-004 选「部分」 → payload:
```
std_ind_001=None, std_ind_002=None, std_ind_003=None, std_ind_004=1000
```
- **CSS 微调**([web/static/style.css:296-299](web/static/style.css#L296-L299)):
新增 `.sample-scope--leaf { margin-left: 8px; }`,比 group 行的 12px 略小,
紧贴 match-inputs 左边缘(match-inputs 自身已有 margin-left: 12px)。
### 9. 自定义规则保持 card-level(一个 step_id 共享一个 limit)
- **决策**:自定义规则保留 card-header 一个 radio(共享 `__custom__` key → 后端
`custom_value_check` step_id),**不**下沉到每条规则。
- **原因**:[web/core/step_impl/step9_custom_rules.py](web/core/step_impl/step9_custom_rules.py)
里所有自定义规则共用一个 `sample_limit`,SQL 构造按 3 路分支(None / oracle / else)
一次性拼接。要下沉到每条规则需要改 step9 接口 + payload 结构(rule-level limit
数组),本次暂不做;用户当前若需 per-rule 控制可以再迭代。
- **UI 一致性折中**:card 头 radio 沿用 `__custom__` key(独立 namespace),不与
analysisTree 的 step_id 撞。
### 提交
待 commit。
---
## 2026-08-13 · Oracle 自定义规则踩坑(ORA-01036 + ORA-00933)+ 多 tab UX 改造
> 接续上面「Oracle 端到端打通」之后的两件事:
......
......@@ -239,6 +239,7 @@ class JobManager:
enable_llm=job.req.enable_llm,
match_overrides=getattr(job.req, "match_overrides", None),
custom_rules=getattr(job.req, "custom_rules", None),
sample_limits=getattr(job.req, "sample_limits", None),
on_log=lambda entry: self._sync_log(job, entry),
on_step_start=lambda step_id, title: self._sync_step_start(job, step_id, title),
on_step_done=lambda step_id, title, ok: self._sync_step_done(job, step_id, title, ok),
......
......@@ -45,6 +45,12 @@ class ConnectRequest(BaseModel):
# SELECT <col> FROM <table> WHERE <col> LIKE '%<keyword>%' ESCAPE '\\' LIMIT 200
# user_input 为空的项会被跳过(不报错)
custom_rules: list[dict] = Field(default_factory=list)
# 可选:每个 step 的「抽样上限」覆盖(2026-08-13 加 partial/full radio 后启用)
# key = step_id(如 "std_ind_001" / "custom_value_check")
# value = int(partial 模式抽样上限,前端默认 1000)
# | None(全量扫描,不加 LIMIT;full 模式)
# 未列出的 step 走 step 自己的默认值(向后兼容)。
sample_limits: dict[str, int | None] = Field(default_factory=dict)
class TestConnectionRequest(ConnectRequest):
......
......@@ -130,7 +130,8 @@ def _run_length_check(*, cfg, dict_data, llm, log, cancel_event, table_filter, m
def _run_standards_one(standard_id: str):
"""返回单个 indicator step 的 runner(捕获 standard_id 闭包)。"""
def _runner(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides):
def _runner(*, cfg, dict_data, llm, log, cancel_event, table_filter,
match_overrides, sample_limit=None):
from .step_impl.step7_standards import run_step7_for_indicator
# step_id 是 std_ind_xxx,但前端 match_overrides 按 standard_id(如 IND-001-a)组织;
# 这里两种 key 都接受,向后兼容。
......@@ -142,6 +143,7 @@ def _run_standards_one(standard_id: str):
return run_step7_for_indicator(
standard_id, cfg, dict_data=dict_data, log=log, table_filter=table_filter,
match_override=override,
sample_limit=sample_limit,
)
return _runner
......@@ -161,9 +163,10 @@ def _run_custom_value_check_factory(custom_rules: list[dict] | None):
dispatch 循环调 runner 时只传 (cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides),
没有 custom_rules 参数;这里通过工厂模式在注册时把 custom_rules 钉死。
"""
def _runner(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides): # noqa: ARG001
def _runner(*, cfg, dict_data, llm, log, cancel_event, table_filter,
match_overrides, sample_limit=None): # noqa: ARG001
from .step_impl.step9_custom_rules import run_step_custom_value_check
data = run_step_custom_value_check(cfg, custom_rules or [], log=log)
data = run_step_custom_value_check(cfg, custom_rules or [], log=log, sample_limit=sample_limit)
return {"section_key": "custom_value_check", "data": data}
return _runner
......@@ -174,10 +177,11 @@ def _run_custom_value_check_factory(custom_rules: list[dict] | None):
_CUSTOM_RULES_CTX: list[dict] | None = None
def _run_custom_value_check(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides): # noqa: ARG001
def _run_custom_value_check(*, cfg, dict_data, llm, log, cancel_event, table_filter,
match_overrides, sample_limit=None): # noqa: ARG001
"""dispatch 调用的 runner —— 读 _CUSTOM_RULES_CTX(workflow 启动时注入)"""
from .step_impl.step9_custom_rules import run_step_custom_value_check
data = run_step_custom_value_check(cfg, _CUSTOM_RULES_CTX or [], log=log)
data = run_step_custom_value_check(cfg, _CUSTOM_RULES_CTX or [], log=log, sample_limit=sample_limit)
return {"section_key": "custom_value_check", "data": data}
......@@ -668,12 +672,14 @@ _register_indicator_steps()
# 共用 _IND_COMBINED 配置 + run_step7_combined 入口。
def _run_standards_combined(parent_id: str):
"""工厂:生成 step 闭包,调用 run_step7_combined(parent_id, ...)。"""
def _runner(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides): # noqa: ARG001
def _runner(*, cfg, dict_data, llm, log, cancel_event, table_filter,
match_overrides, sample_limit=None): # noqa: ARG001
from .step_impl.step7_standards import run_step7_combined
# run_step7_combined 已返回 {section_key, data},直接透传(不要再 wrap)
return run_step7_combined(
parent_id, cfg, dict_data=dict_data, log=log,
table_filter=table_filter, match_overrides=match_overrides,
sample_limit=sample_limit,
)
return _runner
......@@ -805,13 +811,15 @@ for _step_id, _meta in _COMBINED_STEP_META.items():
# 2026-08-13 起:唯一子检查 IND-005-a(固定电话格式)改名为 std_ind_005,
# UI 标题去掉 -a(与 IND-001/002/003/004 视觉一致)。
# step_id 由 indicator_step_id("IND-005-a") = "std_ind_005_a" 改为 "std_ind_005"。
def _run_standards_ind_005(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides): # noqa: ARG001
def _run_standards_ind_005(*, cfg, dict_data, llm, log, cancel_event, table_filter,
match_overrides, sample_limit=None): # noqa: ARG001
from .step_impl.step7_standards import run_step7_for_indicator
# 把它当作普通单 indicator 跑(IND-005-a 没合并到任何复合 step)
# section_key 强制改成 "standards_ind_005",与 step_id 视觉对齐
res = run_step7_for_indicator(
"IND-005-a", cfg, dict_data=dict_data, log=log,
table_filter=table_filter, match_override=match_overrides.get("std_ind_005") if match_overrides else None,
sample_limit=sample_limit,
)
# 端到端重命名:section_key / step_id / tab.key / tab.title
res["section_key"] = "standards_ind_005"
......@@ -852,12 +860,14 @@ register_step(
# 移出 → 进「国标字段规范」,step_id 由 std_ind_006_b 改为 std_ind_006。
# IND-006-a(通讯地址)保留在「业务字段规范」不变(用户只动 b)。
# 与 IND-005-a → IND-005 改名同款模式:section_key / step_id / tab.key / tab.title 4 处同步改。
def _run_standards_ind_006(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides): # noqa: ARG001
def _run_standards_ind_006(*, cfg, dict_data, llm, log, cancel_event, table_filter,
match_overrides, sample_limit=None): # noqa: ARG001
from .step_impl.step7_standards import run_step7_for_indicator
res = run_step7_for_indicator(
"IND-006-b", cfg, dict_data=dict_data, log=log,
table_filter=table_filter,
match_override=match_overrides.get("std_ind_006") if match_overrides else None,
sample_limit=sample_limit,
)
# 端到端重命名:section_key / step_id / tab.key / tab.title
res["section_key"] = "standards_ind_006"
......@@ -915,12 +925,14 @@ def _run_standards_ind_renamed(
`log` 从 orchestrator dispatcher 传进来(不要在闭包里钉死,否则丢日志)
"""
target_section_key = f"standards_{target_step_id[len('std_'):]}" # "standards_ind_017"
def _runner(*, cfg, dict_data, llm, log, cancel_event, table_filter, match_overrides): # noqa: ARG001
def _runner(*, cfg, dict_data, llm, log, cancel_event, table_filter,
match_overrides, sample_limit=None): # noqa: ARG001
from .step_impl.step7_standards import run_step7_for_indicator
res = run_step7_for_indicator(
source_id, cfg, dict_data=dict_data, log=log,
table_filter=table_filter,
match_override=match_overrides.get(target_step_id) if match_overrides else None,
sample_limit=sample_limit,
)
# 端到端重命名:section_key / step_id / tab.key / tab.title
res["section_key"] = target_section_key
......@@ -1063,6 +1075,7 @@ async def run_governance_workflow(
enable_llm: bool = True,
match_overrides: dict[str, dict[str, list[str]]] | None = None,
custom_rules: list[dict] | None = None,
sample_limits: dict[str, int | None] | None = None,
on_log: Callable[[dict], None] | None = None,
on_step_start: Callable[[str, str], None] | None = None,
on_step_done: Callable[[str, str, bool], None] | None = None,
......@@ -1081,6 +1094,11 @@ async def run_governance_workflow(
match_overrides: 每 step 的「字段名 / 字段注释」用户运行时 override
custom_rules: 用户在前端累积的自定义规则(点字段表 + 号 → 输入关键字)
非空时自动追加 custom_value_check 步骤到 main_steps
sample_limits: 每 step 的「抽样上限」覆盖(dict[step_id] -> int|None)
- 正整数:单字段抽样 LIMIT N(partial 模式,前端默认 1000)
- None :不加 LIMIT,全量扫描(full 模式)
未列出的 step 走 step 自己的默认值(向前兼容)。
2026-08-13:前端 analysisTree group 加 partial/full radio;本参数按 step_id 注入。
Returns:
{
......@@ -1250,6 +1268,7 @@ async def run_governance_workflow(
cancel_event=cancel_event,
table_filter=table_filter,
match_overrides=match_overrides,
sample_limit=(sample_limits or {}).get(sd.step_id),
),
)
section_key = output.get("section_key")
......
"""SQL 字符串调整工具(不执行,只做语法级变换)
当前只有 1 个工具:按 sample_limit 调整 LIMIT / ROWNUM 子句
- sample_limit 有值:SQL 保持不变(已经在前面渲染好 LIMIT N 或 ROWNUM <= N)
- sample_limit 为 None:按 db_type 去掉 LIMIT N 或 AND ROWNUM <= N,做全量扫描
为什么不直接在 sql_loader 里加 optional filter:
- 当前 _render_placeholders 只支持「变量缺失」或「变量替换」,没有「条件整段不渲染」
- LIMIT 子句在 MySQL/达梦 是末尾后缀;Oracle 11g 是 WHERE 内子句,语法位置不同
- 调用方在拿到渲染好的 SQL 后做一次后处理最干净,sql_loader 不感知
为什么不放 db_adapter.py:db_adapter 是 DB 驱动层,sql_utils 是 SQL 字符串处理工具,
两者职责不同。like_util 已经按这个模式拆出来了。
"""
from __future__ import annotations
import re
# LIMIT N / LIMIT 0, N(MySQL/达梦)—— 末尾子句
# 注意:只匹配末尾($),不会误伤字段里的 "LIMIT" 字面量
_RE_LIMIT_TRAILING = re.compile(r"\s+LIMIT\s+\d+(?:\s*,\s*\d+)?\s*$", re.IGNORECASE)
# AND ROWNUM <= N(Oracle 11g)—— WHERE 末尾子句
# 只在末尾匹配:sql_builder 拼出来的 Oracle SQL 形如 "WHERE x LIKE :1 ESCAPE '!' AND ROWNUM <= 200"
_RE_ROWNUM_TRAILING = re.compile(
r"\s+AND\s+ROWNUM\s*<=\s*\d+\s*$", re.IGNORECASE,
)
def apply_sample_limit(sql: str, sample_limit: int | None, db_type: str) -> str:
"""按 sample_limit 调整 SQL 的 LIMIT / ROWNUM 子句。
Args:
sql: 已渲染好的 SQL 字符串(来自 sql_loader 或手动拼接)
sample_limit:
- None 或 0 → 不加行数限制(移除现有的 LIMIT / ROWNUM 子句)
- 正整数 → SQL 不变(假设 caller 已经在 SQL 里加了对应数字的子句)
- 但本函数不负责「在 SQL 里加 LIMIT N」—— 那由 caller 自己拼接
(sample_limit 有值时直接调用方传 N 给 sql_loader 即可)
Returns:
调整后的 SQL 字符串。
- sample_limit=None:去掉末尾的 LIMIT N 或 AND ROWNUM <= N
- sample_limit 有值:原样返回(不做校验,caller 应自己保证 SQL 里 LIMIT 跟 sample_limit 一致)
适用 caller:
- step7_standards:渲染 SQL 后调一次(partial=1000→LIMIT 1000;full=None→移除)
- step9_custom_rules:直接拼 SQL(不走 sql_loader),这里只做 full 时的移除兜底
"""
if sample_limit is None:
db = (db_type or "").lower()
if db == "oracle":
sql = _RE_ROWNUM_TRAILING.sub("", sql)
else:
# MySQL / 达梦;其它方言未实现,先按 MySQL 处理
sql = _RE_LIMIT_TRAILING.sub("", sql)
return sql
\ No newline at end of file
......@@ -428,7 +428,7 @@ def _EMPTY_FALLBACK() -> dict:
return _wrap({g: _empty_group_data() for g in _TABS_BY_GROUP.keys()})
SAMPLE_LIMIT = 500 # 每字段最多抽样
SAMPLE_LIMIT = 500 # 每字段最多抽样(默认值,partial 模式由 orchestrator 注入 sample_limit 覆盖)
MAX_TABLES_PER_FIELD = 5 # 每字段最多检查前 N 张表
# 单字段抽样 SQL
......@@ -565,13 +565,23 @@ def _run_round1_one(
std_instance: BaseStandard,
log: Callable | None,
match_override: dict | None = None,
sample_limit: int | None = None,
) -> dict:
"""单值校验:跑一个 indicator 实例(适用于 22 个单值类 indicator)。
返回单 indicator 数据结构(适用 _wrap_single_indicator)。
若该 indicator 在数据字典里没命中任何字段 → 返回 _empty_indicator_data()。
Args:
sample_limit:
- 正整数 → 单字段抽样 LIMIT N(partial 模式)
- None → 不加 LIMIT,全量扫描(full 模式)
- 默认 None 时由 caller 用 SAMPLE_LIMIT 兜底
"""
loader = get_sql_loader()
from ..sql_utils import apply_sample_limit
# 实际渲染 SQL 时用的 limit:有 sample_limit 用 sample_limit;None 用默认值兜底
effective_limit = sample_limit if sample_limit is not None else SAMPLE_LIMIT
bucket = _empty_indicator_data()
std_id = std_instance.standard_id
......@@ -640,8 +650,10 @@ def _run_round1_one(
dialect=cfg.db_type,
table=table,
field=field_name,
limit=SAMPLE_LIMIT,
limit=effective_limit,
)
# 2026-08-13:sample_limit=None → 移除末尾 LIMIT N(full 模式)
sql = apply_sample_limit(sql, sample_limit, cfg.db_type)
rows = db.fetchall(sql)
sample_sql_rendered = sql
except Exception as e:
......@@ -807,11 +819,18 @@ def _run_round2_cross_field_one(
cfg: DBConfig,
by_name: dict[str, list[dict]],
log: Callable | None,
sample_limit: int | None = None,
) -> dict:
"""跨字段校验 IND-301:证件类型 + 证件号 一致性 —— 单 indicator 输出。"""
"""跨字段校验 IND-301:证件类型 + 证件号 一致性 —— 单 indicator 输出。
Args:
sample_limit: 同 _run_round1_one;None 表示全量扫描
"""
from standards.ind_301_id_type_number_consistency import IdTypeNumberConsistencyIndicator
loader = get_sql_loader()
from ..sql_utils import apply_sample_limit
effective_limit = sample_limit if sample_limit is not None else SAMPLE_LIMIT
bucket = _empty_indicator_data()
ind301 = IdTypeNumberConsistencyIndicator()
......@@ -861,8 +880,10 @@ def _run_round2_cross_field_one(
table=table,
type_col=tc,
no_col=nc,
limit=SAMPLE_LIMIT,
limit=effective_limit,
)
# 2026-08-13:sample_limit=None → 移除末尾 LIMIT N(full 模式)
sql = apply_sample_limit(sql, sample_limit, cfg.db_type)
rows = db.fetchall(sql)
except Exception as e:
if log:
......@@ -944,8 +965,13 @@ def _run_round2_uniqueness_one(
cfg: DBConfig,
by_name: dict[str, list[dict]],
log: Callable | None,
sample_limit: int | None = None, # noqa: ARG001 # 保留签名一致;本函数内部走 GROUP BY HAVING,不需要 sample
) -> dict:
"""跨表聚合 IND-302:证件唯一性 GROUP BY HAVING —— 单 indicator 输出。"""
"""跨表聚合 IND-302:证件唯一性 GROUP BY HAVING —— 单 indicator 输出。
sample_limit 不影响本函数(GROUP BY HAVING 本身就是全表聚合,不需要单字段抽样)。
保留参数是为了让 caller 入口签名一致,简化 dispatch。
"""
from standards.ind_302_id_uniqueness import IdUniquenessIndicator
loader = get_sql_loader()
......@@ -1081,11 +1107,14 @@ def _run_round3_global_uniqueness_one(
cfg: DBConfig,
by_name: dict[str, list[dict]],
log: Callable | None,
sample_limit: int | None = None, # noqa: ARG001 # 保留签名一致;本函数走 GROUP BY HAVING 不需要 sample
) -> dict:
"""跨表单字段 GROUP BY HAVING —— IND-901 单位名称全局唯一性。
按字段名白名单 + 注释命中 找单位名称字段,按表跑 GROUP BY 单字段 SQL,
检出同表重复值。
sample_limit 不影响本函数(GROUP BY HAVING 本身就是全表聚合)。
"""
from standards.ind_901_company_name_global_uniqueness import (
CompanyNameGlobalUniquenessIndicator,
......@@ -1225,6 +1254,7 @@ def _run_round3_cross_field_one(
cfg: DBConfig,
by_name: dict[str, list[dict]],
log: Callable | None,
sample_limit: int | None = None,
) -> dict:
"""跨字段一致性 —— IND-902 账户名 vs 参考字段(同行配对)。
......@@ -1232,7 +1262,12 @@ def _run_round3_cross_field_one(
- 账户名侧:account_name / acct_name / yhmc / 收款账户名 等
- 参考侧:单位名称 / 个人姓名 / 法人姓名 等
按表跑 sample_field_pairs_two_text.sql,调 validate_rows 比对一致性。
Args:
sample_limit: 同 _run_round1_one;None 表示全量扫描
"""
from ..sql_utils import apply_sample_limit
effective_limit = sample_limit if sample_limit is not None else SAMPLE_LIMIT
from standards.ind_902_account_name_consistency import (
AccountNameConsistencyIndicator,
)
......@@ -1304,8 +1339,10 @@ def _run_round3_cross_field_one(
table=table,
col_a=ac,
col_b=bc,
limit=SAMPLE_LIMIT,
limit=effective_limit,
)
# 2026-08-13:sample_limit=None → 移除末尾 LIMIT N(full 模式)
sql = apply_sample_limit(sql, sample_limit, cfg.db_type)
rows = db.fetchall(sql)
except Exception as e:
if log:
......@@ -1485,12 +1522,18 @@ def run_step7_for_indicator(
log: Callable | None = None,
table_filter: set[str] | None = None,
match_override: dict | None = None,
sample_limit: int | None = None,
) -> dict:
"""跑单个 indicator(orchestrator 每个 step 调一次)。
Args:
standard_id: 例如 "IND-001-a" / "IND-301" / "IND-302"
cfg / dict_data / log / table_filter: 同 run_step7
match_override: 用户 UI 改的 applies_to_fields / comment_keywords 覆盖
sample_limit:
- 正整数:单字段抽样 LIMIT N(partial 模式,前端默认 1000)
- None :不加 LIMIT,全量扫描(full 模式)
2026-08-13:前端 analysisTree group 加 partial/full radio;本参数由 orchestrator 按 step_id 注入
Returns:
{section_key, data} —— data 已注入 _protocol(1 tab)
......@@ -1521,13 +1564,13 @@ def run_step7_for_indicator(
by_name.setdefault(r["column_name"], []).append(r)
if standard_id == "IND-301":
data = _run_round2_cross_field_one(cfg, by_name, log)
data = _run_round2_cross_field_one(cfg, by_name, log, sample_limit=sample_limit)
elif standard_id == "IND-302":
data = _run_round2_uniqueness_one(cfg, by_name, log)
data = _run_round2_uniqueness_one(cfg, by_name, log, sample_limit=sample_limit)
elif standard_id == "IND-901":
data = _run_round3_global_uniqueness_one(cfg, by_name, log)
data = _run_round3_global_uniqueness_one(cfg, by_name, log, sample_limit=sample_limit)
elif standard_id == "IND-902":
data = _run_round3_cross_field_one(cfg, by_name, log)
data = _run_round3_cross_field_one(cfg, by_name, log, sample_limit=sample_limit)
else:
# 单值类:字段名 OR 注释命中都可(applies_to_fields / comment_keywords 二选一非空)
# 用 _effective_match 合并 YAML + 用户 override 后再判定
......@@ -1542,6 +1585,7 @@ def run_step7_for_indicator(
data = _run_round1_one(
cfg, columns, std_instance, log,
match_override=match_override,
sample_limit=sample_limit,
)
return {"section_key": section_key, "data": _wrap_single_indicator(std_instance, data)}
......@@ -1821,6 +1865,7 @@ def run_step7_combined(
log: Callable | None = None,
table_filter: set[str] | None = None,
match_overrides: dict[str, dict] | None = None,
sample_limit: int | None = None,
) -> dict:
"""合并子检查的入口:跑 parent_id 下全部子检查,合并为 1 section / 1 tab。
......@@ -1831,6 +1876,7 @@ def run_step7_combined(
- "std_ind_001" 合并 step 的整体 override(推荐)
- "IND-001-a" / "-b" / "-c" / "-d" 单子检查 override(向后兼容老 payload)
- "std_ind_001_a" 等 orchestrator 风格的 step_id key(向后兼容)
sample_limit: 同 run_step7_for_indicator;本函数把 limit 透传给所有子检查的 _run_round1_one
Returns:
{"section_key": "standards_ind_001", "data": <merged _protocol>}
......@@ -1898,6 +1944,7 @@ def run_step7_combined(
single = _run_round1_one(
cfg, columns, std, log,
match_override=override,
sample_limit=sample_limit,
)
combined["violations_by_indicator"].extend(single["violations_by_indicator"])
combined["violations"].extend(single["violations"])
......
......@@ -61,7 +61,7 @@ from ..tab_protocol import (
logger = logging.getLogger(__name__)
_LIMIT = 200 # SQL 单条查询上限(不做 distinct,直接按行输出)
_DEFAULT_LIMIT = 200 # SQL 单条查询默认值(partial 模式由 orchestrator 注入 sample_limit 覆盖)
# ── payload 兼容:旧 {rule_type, user_input} → 新 {group: ...} ──
......@@ -93,6 +93,7 @@ def run_step_custom_value_check(
cfg: DBConfig,
custom_rules: list[dict],
log: Callable | None = None,
sample_limit: int | None = None,
) -> dict:
"""对每条 user 规则的「条件树」编译成 WHERE → 跑 SQL → 写入 matches[] + _protocol。
......@@ -102,6 +103,11 @@ def run_step_custom_value_check(
group: {kind:'group', op, children:[...]}, ...}]
或旧 payload [{rule_type, user_input, ...}]
log: 日志回调(orchestrator 注入)
sample_limit: 单条规则 SQL 的命中行数上限
- 正整数:加 LIMIT N(MySQL/达梦)或 AND ROWNUM <= N(Oracle 11g)
- None: 不加行数限制,全量扫描
- 默认 None(旧调用方不传时由 orchestrator 用 _DEFAULT_LIMIT=200 兜底)
2026-08-13:前端分析项目加 partial/full radio;选 full → None
Returns:
{
......@@ -208,22 +214,31 @@ def run_step_custom_value_check(
skipped_no_valid_leaf += 1
continue
# 截断子句按 db_type 分支:MySQL/达梦走 LIMIT;Oracle 11g 不支持 LIMIT/FETCH,
# 截断子句按 db_type + sample_limit 分支:
# - sample_limit 有值:MySQL/达梦走 LIMIT N;Oracle 11g 不支持 LIMIT/FETCH,
# 用 WHERE ROWNUM <= N(兼容 11.2.0.4;12c+ 也仍支持 ROWNUM 写法)。
# 2026-08-13:Oracle 上跑自定义规则报 ORA-00933 暴露此差异。
# where_clause 已由 L187 保证非空,所以可直接拼 AND ROWNUM <= N。
if cfg.db_type == "oracle":
# - sample_limit=None:不加任何行数限制,全量扫描。
# 2026-08-13:Oracle 上跑自定义规则报 ORA-00933 暴露方言差异。
# 2026-08-13:前端加 partial/full radio;full → sample_limit=None → 不加 LIMIT。
# where_clause 已由 L187 保证非空,所以 partial + Oracle 可直接拼 AND ROWNUM <= N。
if sample_limit is None:
sql = (
f"SELECT {col_quoted} "
f"FROM {quote_ident(table_name, cfg.db_type)} "
f"WHERE {where_clause} AND ROWNUM <= {_LIMIT}"
f"WHERE {where_clause}"
)
elif cfg.db_type == "oracle":
sql = (
f"SELECT {col_quoted} "
f"FROM {quote_ident(table_name, cfg.db_type)} "
f"WHERE {where_clause} AND ROWNUM <= {sample_limit}"
)
else:
sql = (
f"SELECT {col_quoted} "
f"FROM {quote_ident(table_name, cfg.db_type)} "
f"WHERE {where_clause} "
f"LIMIT {_LIMIT}"
f"LIMIT {sample_limit}"
)
try:
......
......@@ -445,6 +445,18 @@
<el-tag v-if="row.llm_mode === 'required'" size="small" type="warning" effect="plain" class="badge">需 LLM</el-tag>
<el-tag v-if="row.required" size="small" type="danger" effect="dark" class="badge">必选</el-tag>
</el-checkbox>
<!-- 2026-08-13:抽样范围 partial/full radio(per-IND)
key 用 step_id(与后端 sample_limits[step_id] 一致)
位置在 match-inputs(字段名输入框)左侧 -->
<el-radio-group
class="sample-scope sample-scope--leaf"
v-model="sampleScope[row.step_id]"
size="small"
@click.stop
>
<el-radio-button label="partial">部分</el-radio-button>
<el-radio-button label="full">全量</el-radio-button>
</el-radio-group>
<!-- 字段匹配输入框:仅在 step 命中 YAML 配置时显示;跨字段/跨表 indicator 不显示 -->
<span
v-if="row.step_id && matchConfig[row.step_id] && !matchConfig[row.step_id].skip"
......@@ -537,6 +549,17 @@
<el-icon><Operation /></el-icon>
<span class="card-title">自定义规则</span>
<span class="dict-status">已添加 <strong style="color: #409eff">{{ customRules.length }}</strong> 项</span>
<!-- 2026-08-13:自定义规则抽样范围 partial/full radio
自定义规则不在 analysisTree 的 group 里,独立占一个 key "__custom__" -->
<el-radio-group
class="sample-scope"
v-model="sampleScope['__custom__']"
size="small"
@click.stop
>
<el-radio-button label="partial">部分</el-radio-button>
<el-radio-button label="full">全量</el-radio-button>
</el-radio-group>
<div style="margin-left: auto">
<el-button size="small" @click.stop="clearCustomRules" :disabled="customRules.length === 0">清空</el-button>
<el-button
......@@ -1157,8 +1180,22 @@
// analysisTree 加载后不再自动展开任何组;保留该 watcher 占位
// 以备未来需要在此做初始化(例如按 group key 推断默认勾选)。
// 2026-08-13:顺便初始化 sampleScope 默认值 'partial',避免 v-model 绑定 undefined。
// per-IND radio 改成 key=step_id(每项独立配置),不再按 group 共享。
watch(() => analysisTree.value.groups, (_groups) => {
/* 默认全部折叠 —— 不做任何自动展开 */
if (Array.isArray(_groups)) {
for (const g of _groups) {
for (const c of (g.checks || [])) {
if (c && c.step_id && !sampleScope.value[c.step_id]) {
sampleScope.value[c.step_id] = 'partial';
}
}
}
}
if (!sampleScope.value['__custom__']) {
sampleScope.value['__custom__'] = 'partial';
}
}, { immediate: true });
// ── 数据字典(来自 connect/test 响应) ──
......@@ -1195,6 +1232,16 @@
// rule_type 当前固定为 'contains_keyword'(占位,后续会扩展更多规则类型)
const customRules = ref([]);
const customConfigCollapsed = ref(false);
// ── 抽样范围(per-IND:partial / full)──
// 2026-08-13:用户反馈「针对特定字段的分析希望可切换部分/全量」
// - key = step_id(如 "std_ind_001" / "length_check")
// 或自定义规则的固定 key "__custom__"
// - value = 'partial'(默认,前 N 行抽样,partial_limit=1000)
// | 'full'(全量,不加 LIMIT)
// analysisTree 加载后 watcher 会为每个 step_id 初始化 'partial',
// 避免 v-model 绑定 undefined 导致 radio 不显示选中态。
const sampleScope = ref({});
const SAMPLE_PARTIAL_LIMIT = 1000; // 前端默认 partial 上限(与后端约定)
const logs = ref([]);
const logExpanded = ref(false); // 2026-08-12 改:默认折叠实时日志
const showStandardsDialog = ref(false);
......@@ -1848,6 +1895,9 @@
match_overrides: buildMatchOverridesForSubmit(),
// 自定义规则:保留 user_input 非空的项;空关键字在后端会被静默跳过
custom_rules: buildCustomRulesForSubmit(),
// 抽样范围:partial(默认 1000)/ full(无 LIMIT)
// 后端 models.sample_limits: dict[step_id, int|None]
sample_limits: buildSampleLimitsForSubmit(),
};
try {
const r = await fetch('/api/jobs', {
......@@ -1937,6 +1987,30 @@
return false;
}
// ── 提交时把 sampleScope 转成后端要的 {step_id: int|None} 字典 ──
// 输入:sampleScope.value = { '<step_id>': 'partial'|'full' (per-IND radio)
// '__custom__': 'partial'|'full' (自定义规则 card 共享) }
// 输出:{ '<step_id>': <int>|None, ... }
// - 'full' → None (后端会移除 LIMIT N / AND ROWNUM<=N)
// - 'partial' → SAMPLE_PARTIAL_LIMIT (1000)
// - 缺省值 → 'partial'(默认部分;与 UI 默认一致)
// 自定义规则固定映射到 step_id='custom_value_check'(后端 step9 标识;
// 同一个 step_id 共享一个 sample_limit —— step9 内部对所有规则用同一 LIMIT)
function buildSampleLimitsForSubmit() {
const out = {};
const groups = (analysisTree.value && analysisTree.value.groups) || [];
for (const g of groups) {
for (const c of (g.checks || [])) {
const scope = sampleScope.value[c.step_id];
out[c.step_id] = scope === 'full' ? null : SAMPLE_PARTIAL_LIMIT;
}
}
// 自定义规则(独立卡片,不在 analysisTree 的 group 里;card 内多规则共享一个 limit)
const customScope = sampleScope.value['__custom__'];
out['custom_value_check'] = customScope === 'full' ? null : SAMPLE_PARTIAL_LIMIT;
return out;
}
// ── 订阅日志 ──
function subscribeLogs(jobId) {
if (logEventSource) logEventSource.close();
......@@ -2243,6 +2317,7 @@
analysisConfigCollapsed, dataDictCollapsed, connectionCollapsed,
customRules, customConfigCollapsed,
addCustomRule, removeCustomRule, clearCustomRules,
sampleScope, SAMPLE_PARTIAL_LIMIT, buildSampleLimitsForSubmit,
fieldClassTagType,
onDbTypeChange,
selectedField, selectFieldFilter, clearFieldFilter,
......
......@@ -278,6 +278,36 @@ body {
min-width: 0; /* 让 ellipsis 在 flex 容器中生效 */
flex: 1; /* 占据剩余宽度 */
}
/* 2026-08-13:抽样范围 partial/full radio
- 限制 description 最大宽度,给 radio 让位(避免换行错位)
- radio 自身在行内最右,flex-shrink:0 防被挤压 */
.tree-row--group .group-description {
max-width: 280px;
flex: 0 1 auto;
}
.sample-scope {
margin-left: 12px;
flex-shrink: 0;
}
.sample-scope .el-radio-button__inner {
padding: 5px 10px;
font-size: 12px;
}
/* 叶子行(per-IND)的 radio:与 match-inputs 之间紧贴,左边距小一些 */
.sample-scope--leaf {
margin-left: 8px;
}
/* 自定义规则卡片里也加一份 radio(在 card 顶部一行右对齐) */
.custom-rule-sample {
display: flex;
align-items: center;
gap: 8px;
margin-bottom: 8px;
}
.custom-rule-sample .label {
color: #606266;
font-size: 13px;
}
.tree-row .badge {
margin-left: 6px;
height: 18px;
......
Markdown is supported
0%
or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment