构建多模态异常检测数据集MMR-AD,推动大模型在工业异常检测中的应用
MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
- 构建包含图像、文本的多模态数据集MMR-AD,专为训练和评估多模态大模型设计
- 现有主流大模型在异常检测任务上表现远低于工业需求,基准测试显示差距明显
- 提出基于思维链推理与强化学习的Anomaly-R1模型,显著提升检测与定位能力
在工业异常检测的发展中,通用异常检测(GAD)已成为新兴趋势和最终目标。与传统的单类或多元类异常检测不同,通用异常检测旨在训练一个无需在目标数据上重新训练或微调即可直接检测多种新类别异常的通用模型。近年来,多模态大语言模型(MLLMs)因其强大的视觉理解与语言推理能力,在实现通用异常检测方面展现出巨大潜力。然而,由于(1)当前主流的MLLM预训练数据主要来源于网络,与真实工业异常检测场景存在显著差距,且预训练中的图文对并非专为异常检测任务设计;(2)现有主流异常检测数据集均为图像驱动,难以支持后续对MLLM的微调。为此,我们提出了MMR-AD——一个面向训练和评估基于多模态大模型的异常检测模型的综合性基准数据集。基于该数据集,我们发现当前最先进通用型大模型在异常检测性能上仍远未达到工业要求。在此基础上,我们进一步提出基线模型Anomaly-R1,该模型通过学习MMR-AD中的思维链(CoT)数据,并结合强化学习进行优化,在异常检测与定位任务上均显著优于现有通用大模型。
原文摘要 · Abstract (English)
In the progress of industrial anomaly detection, general anomaly detection (GAD) is an emerging trend and also the ultimate goal. Unlike the conventional single- and multi-class AD, general AD aims to train a general AD model that can directly detect anomalies in diverse novel classes without any retraining or fine-tuning on the target data. Recently, Multimodal Large Language Models (MLLMs) have shown great promise in achieving general anomaly detection due to their revolutionary visual understanding and language reasoning capabilities. However, MLLM's general AD ability remains underexplored due to: (1) MLLMs are pretrained on amounts of data sourced from the Web, these data still have significant gaps with the data in AD scenarios. Moreover, the image-text pairs during pretraining are also not specifically for AD tasks. (2) The current mainstream AD datasets are image-based and not yet suitable for post-training MLLMs. To facilitate MLLM-based general AD research, we present MMR-AD, which is a comprehensive benchmark for both training and evaluating MLLM-based AD models. With MMR-AD, we reveal that the AD performance of current SOTA generalist MLLMs still falls far behind the industrial requirements. Based on MMR-AD, we also propose a baseline model, Anomaly-R1, which is a reasoning-based AD model that learns from the CoT data in MMR-AD and is further enhanced by reinforcement learning. Extensive experiments show that our Anomaly-R1 achieves remarkable improvements over generalist MLLMs in both anomaly detection and localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。