arXiv:2607.16243cs.LGcs.AI2026-07中稿 · publication in Tra…

首个面向工业部署的多模态小模型鲁棒性评测基准

RobustMAD: Evaluating Real-World Robustness of Multimodal Small Language Models for Deployable Anomaly Detection Assistants

论文配图:RobustMAD: Evaluating Real-World Robustness of Multimodal Small Language Models for Deployable Anomaly Detection Assistants
图 1 · 摘自论文原文
  • 构建真实工业场景下的多模态评测基准,覆盖多种视觉退化与开放问题
  • 小模型在部分任务上超越大模型,但存在三类致命缺陷
  • 适合关注工业AI落地安全性的研发与评估人员

多模态工业异常检测助手是下一代智能工厂的核心,支持基于视觉与语言的交互式查询。然而,多模态大模型因计算成本高和云端推理带来的隐私风险,难以现场部署。紧凑型多模态小语言模型(MSLMs)提供了可部署的替代方案,但其发展受限于缺乏全面的鲁棒性分析和反映真实工业条件的挑战性评测基准。为此,我们提出了RobustMAD,首个以部署为导向的评测基准,通过涵盖物体理解、异常检测、无解问题及视觉质量退化的多样化开放问题,全面评估模型鲁棒性。与传统假设相反,表现最佳的MSLMs展现出令人意外的能力,甚至优于更大的GPT-5 Nano。然而,它们仍无法满足关键安全要求,RobustMAD揭示了严重的鲁棒性差距,带来实际运营风险。具体表现为三种反复出现的失败模式:(i) 在细粒度区分或视觉退化条件下多模态对齐脆弱;(ii) 回答不完整;(iii) 对无解或表述不清的问题逻辑对齐弱,导致幻觉输出。基于这些发现,我们为下一代多模态工业检测助手的设计提供可操作建议。代码已开源。

原文摘要 · Abstract (English)

Multimodal industrial anomaly inspection assistants are a critical component of next-generation smart factories, enabling interactive vision-language-based querying. However, multimodal large language models remain impractical for on-site deployment due to prohibitive computational demands and privacy risks from cloud-based inference. Compact multimodal small language models (MSLMs) offer a deployable alternative, yet progress is constrained by the lack of comprehensive robustness analyses and meaningfully challenging benchmarks that reflect real-world industrial conditions. To address this gap, we develop RobustMAD, the first deployment-motivated benchmark, designed to comprehensively evaluate model robustness through diverse open-ended queries spanning object understanding, anomaly detection, unanswerable problems, and visual quality degradations. Contrary to conventional assumptions, top-performing MSLMs exhibit promising capabilities, surprisingly outperforming even the larger GPT-5 Nano. However, they still fall short of safety-critical requirements, and RobustMAD reveals critical robustness gaps that pose operational risks. In particular, three recurring failure modes emerge: (i) fragile multimodal grounding under fine-grained distinctions or degraded visual conditions, (ii) insufficiently comprehensive responses, and (iii) weak logical grounding on unanswerable or ill-posed queries, leading to hallucinated outputs. Grounded in these insights, we provide actionable guidance for the design of next-generation multimodal industrial inspection assistants that leverage their promising competence. Code is available at https://github.com/en-research/RobustMAD.

多模态小模型工业检测鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。