用6690道题测试大模型从工业规则推荐维修动作,发现模型易被干扰误导。
DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules

- 将规则转为多选题,用嵌入抽样生成干扰项,构建标准化评测流程。
- 顶级模型准确率仅45%,在条件变化下性能下降13%~60%,暴露脆弱性。
- 适合研究工业AI决策可靠性的学者,尤其关注模型鲁棒性与部署风险。
复杂工业设备的监测依赖工程师编写的符号化规则,根据传感器状态触发并提示技术人员执行维护操作。瓶颈不在于故障检测,而在于响应:将规则转化为具体维护步骤需要多年积累的设备知识。我们探究大模型能否作为此类规则到动作转换的决策支持工具,并提出 extbf{DiagnosticIQ},一个包含6,690个专家验证的多选题数据集,源自16类设备上的118组规则-动作对。我们贡献了:(i) 一种符号规则转多选题的流水线,将规则归一化为析取范式,并使用嵌入方法采样干扰项;(ii) 五种变体用于探测不同失效模式(Pro、Pert、Verbose、Aug、Rationale);(iii) 包含29个大模型和4个嵌入基线的评测基准。人工评估(9名从业者,平均准确率45.0%)证实,该任务需要超越日常经验的专业知识。三个发现突出:顶尖模型差距已缩小,前三名仅差1个宏观分数点,但所有模型在干扰项扩展下均出现13%–60%相对准确率下降;在条件反转场景中,前沿模型仍以49%–63%概率选择原答案,暴露模式匹配倾向。部署瓶颈并非能力不足,而是校准问题:模型能应对模板式故障检测,但在结构扰动下失效。
原文摘要 · Abstract (English)
Monitoring complex industrial assets relies on engineer-authored symbolic rules that trigger based on sensor conditions and prompt technicians to perform corrective actions. The bottleneck is not detection but response: translating rules into maintenance steps requires asset-specific knowledge gained through years of practice. We investigate whether LLMs can serve as decision support for this rule-to-action step and introduce \ours{}, a benchmark of 6{,}690 expert-validated multiple-choice questions from 118 rule-action pairs across 16 asset types. We contribute (i) a symbolic-to-MCQA pipeline normalizing rules to Disjunctive Normal Form with embedding-based distractor sampling, (ii) five variants probing distinct failure modes (Pro, Pert, Verbose, Aug, Rationale), and (iii) a benchmark of 29 LLMs and 4 embedding baselines. A human evaluation (9 practitioners, mean 45.0\%) confirms \ours{} requires specialist knowledge beyond operational experience. Three findings stand out. The frontier has closed: the top three LLMs lie within one Macro point, with Bradley-Terry Elo placing claude-opus-4-6 30 points above the next model. Yet \ours{}\,Pro exposes brittleness, with every model losing 13--60\% relative accuracy under distractor expansion. \ours{}\,Aug exposes pattern-matching: under condition inversion, frontier models still select the original answer 49--63\% of the time. The deployment bottleneck is not capability but calibration: frontier models handle template-style fault detection but break under structural perturbation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。