arXiv:2603.00055cs.LGcs.AI2026-03

提出可自我反思的多模态工业缺陷检测框架,提升模型判断可靠性

M3-AD: Reflection-aware Multi-modal, Multi-category, and Multi-dimensional Benchmark and Framework for Industrial Anomaly Detection

  • 构建双数据集:用于反思微调的M3-AD-FT和跨类别评估的M3-AD-Bench
  • 引入可学习的反思机制,让模型在误判时主动修正决策,显著提升鲁棒性
  • 适用于需要高可靠性的工业质检场景,尤其适合零样本缺陷识别任务

尽管多模态大语言模型(MLLMs)已推动工业缺陷检测迈向零样本范式,但在细粒度、结构复杂的工业场景中仍易产生高置信度但不可靠的判断,缺乏有效自纠错机制。为此,我们提出M3-AD,一个统一的、具备反思能力的多模态工业缺陷检测框架。M3-AD包含两个互补的数据资源:用于反思对齐微调的M3-AD-FT,以及用于系统化跨类别评估的M3-AD-Bench,共同为反思学习与可靠性评估提供基础。在此基础上,我们提出RA-Monitor,将反思建模为可学习的决策修订过程,在初始判断不可靠时引导模型进行可控自纠正,从而增强决策鲁棒性。在M3-AD-Bench上的大量实验表明,RA-Monitor在零样本缺陷检测与分析任务中优于多个开源及商用MLLMs。代码将发布于https://github.com/Yanhui-Lee/M3-AD。

原文摘要 · Abstract (English)

Although multimodal large language models (MLLMs) have advanced industrial anomaly detection toward a zero-shot paradigm, they still tend to produce high-confidence yet unreliable decisions in fine-grained and structurally complex industrial scenarios, and lack effective self-corrective mechanisms. To address this issue, we propose M3-AD, a unified reflection-aware multimodal framework for industrial anomaly detection. M3-AD comprises two complementary data resources: M3-AD-FT, designed for reflection-aligned fine-tuning, and M3-AD-Bench, designed for systematic cross-category evaluation, together providing a foundation for reflection-aware learning and reliability assessment. Building upon this foundation, we propose RA-Monitor, which models reflection as a learnable decision revision process and guides models to perform controlled self-correction when initial judgments are unreliable, thereby improving decision robustness. Extensive experiments conducted on M3-AD-Bench demonstrate that RA-Monitor outperforms multiple open-source and commercial MLLMs in zero-shot anomaly detection and anomaly analysis tasks. Code will be released at https://github.com/Yanhui-Lee/M3-AD.

工业质检多模态自我反思零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。