Commit fbedfe12 authored by Data Governance Dev's avatar Data Governance Dev

fix(llm): Step 2 大量 llm_failed 解析失败 — max_tokens 1024 截断大批量响应

症状:最近 run 里 Step 2「表合并与冗余字段分析」21/21 高频字段全部
classification=unknown, source=llm_failed。但 LLM 调用本身成功
(HTTP 200),只是解析阶段拿不到 valid JSON。

根因:web/configs/llm.yaml 默认 max_tokens=1024。Step 2 让 LLM 返回
每条 reasoning+recommendation 各 50-100 字中文 ≈ 200-300 tokens,
15 字段 × ~250 tokens ≈ 3750 tokens 远超 1024 上限 → LLM 响应被截断
在某个汉字中间 → JSON 不完整 → _safe_parse_json 返回 None → 21 个
字段全标 unknown。

复现:用真实 15 字段 prompt 调 complete(),返回 2824 字符但末尾
"\u5e76\u4e14\u5bb9\u6613\u88ab'" 明显是截断。

修复(A+B 一起):

A. web/configs/llm.yaml:max_tokens 1024 → 4096,附注释说明为什么
B. web/core/step_impl/step2_merge_redundancy.py:BATCH 30 → 8,
   附注释说明为什么

实测验证(重启后):8 字段 batch 全部解析成功,分类到 4 种
(suspicious/common_business/common_base/true_redundancy),没有
任何 None 失败。

下次跑 21 字段 → 3 个 batch(8+8+5),每个都能塞进 4096 tokens。

注意:需重启后端服务才能让两处都生效(llm.py 是 Python 代码,
llm.yaml 是启动时 LLMConfig.load() 一次缓存到 _client 单例)。
parent e6fbfb51
...@@ -236,7 +236,7 @@ def _find_redundancy(by_table: dict, llm, log: Callable | None) -> list[dict]: ...@@ -236,7 +236,7 @@ def _find_redundancy(by_table: dict, llm, log: Callable | None) -> list[dict]:
} }
for f, n in candidates for f, n in candidates
] ]
BATCH = 30 BATCH = 8 # 每批 LLM 调用最多 8 个字段(reasoning+recommendation 单条约 200-300 tokens,整批控制在 max_tokens=4096 内)
annotated: list[dict | None] = [] annotated: list[dict | None] = []
total_batches = (len(llm_inputs) + BATCH - 1) // BATCH total_batches = (len(llm_inputs) + BATCH - 1) // BATCH
for batch_idx in range(0, len(llm_inputs), BATCH): for batch_idx in range(0, len(llm_inputs), BATCH):
......
Markdown is supported
0%
or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment