arXiv:2605.10267cs.AI2026-05被引 2

测试大模型在工业采购中的真实可靠性,发现其答案常藏安全漏洞。

IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs

论文配图:IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
图 1 · 摘自论文原文
  • 基于中国国标构建2049项工业采购问答,分7维度10行业
  • 顶尖模型正确率仅2.083/3,70.3%生成答案被拒
  • 长推理反而增加安全风险,需源文本验证

在工业采购中,大模型的回答只有通过标准检验才有用:推荐材料必须匹配工况,参数不能超限,流程不得违背安全条款。部分正确可能掩盖关键矛盾,现有基准难以捕捉此类问题。我们提出IndustryBench,一个包含2049个中文工业采购问答的基准,依据中国国家标准(GB/T)和结构化产品记录构建,涵盖7个能力维度、10个行业类别及专家评定的难度层级,并提供英文、俄文、越南文对齐版本。构建流程在基于搜索的外部验证阶段淘汰了70.3%的模型生成候选答案,揭示仅靠大模型过滤仍不可靠。评估分离原始正确性(由经验证的Qwen3-Max判官评分,κ_w=0.798)与独立的安全违规(SV)检查。在17个中文模型及4语言交叉的8模型中发现:(i) 最佳系统得分仅2.083/3,仍有巨大提升空间;(ii) 标准与术语理解是最顽固的能力短板,且在翻译后依然存在;(iii) 延伸推理使12/13模型的安全调整分数下降,主因是长回答引入无支持的安全关键错误;(iv) 安全违规率重排排名——GPT-5.4经调整后从第6升至第3,Kimi-k2.5-1T-A32B则下滑7位。工业大模型评估需基于源文本、关注安全的诊断,而非单一准确率。数据集及全部提示、评分脚本已开源。

原文摘要 · Abstract (English)

In industrial procurement, an LLM answer is useful only if it survives a standards check: recommended material must match operating condition, every parameter must respect a regulated threshold, and no procedure may contradict a safety clause. Partial correctness can mask safety-critical contradictions that aggregate LLM benchmarks rarely capture. We introduce IndustryBench, a 2,049-item benchmark for industrial procurement QA in Chinese, grounded in Chinese national standards (GB/T) and structured industrial product records, organized by seven capability dimensions, ten industry categories, and panel-derived difficulty tiers, with item-aligned English, Russian, and Vietnamese renderings. Our construction pipeline rejects 70.3% of LLM-generated candidates at a search-based external-verification stage, calibrating how unreliable industrial QA remains after LLM-only filtering. Our evaluation decouples raw correctness, scored by a Qwen3-Max judge validated at $κ_w = 0.798$ against a domain expert, from a separate safety-violation (SV) check against source texts. Across 17 models in Chinese and an 8-model intersection over four languages, we find: (i) the best system reaches only 2.083 on the 0--3 rubric, leaving substantial headroom; (ii) Standards & Terminology is the most persistent capability weakness and survives item-aligned translation; (iii) extended reasoning lowers safety-adjusted scores for 12 of 13 models, primarily by introducing unsupported safety-critical details into longer final answers; and (iv) safety-violation rates reshuffle the leaderboard -- GPT-5.4 climbs from rank 6 to rank 3 after SV adjustment, while Kimi-k2.5-1T-A32B drops seven positions. Industrial LLM evaluation therefore requires source-grounded, safety-aware diagnosis rather than aggregate accuracy. We release IndustryBench with all prompts, scoring scripts, and dataset documentation.

工业AI大模型评测安全验证多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。