用大模型+强化学习实现工业缺陷检测的端到端自动分析
AnomalyR1: A GRPO-based End-to-end MLLM for Industrial Anomaly Detection
- 结合视觉语言模型与强化学习,端到端处理图像和领域知识
- 30亿参数模型在最新基准上达到顶尖性能,优于现有方法
- 适合缺乏缺陷样本的工业场景,可解释性强,便于部署
工业异常检测(IAD)因缺陷样本稀缺而面临严峻挑战,亟需具备强泛化能力的模型以有效识别未见异常。传统方法受限于手工特征或领域专家模型,难以突破此瓶颈,亟需范式革新。本文提出AnomalyR1,首个基于多模态大语言模型(MLLM)VLM-R1的端到端IAD框架,融合组相对策略优化(GRPO)与新颖的有理结果对齐度量(ROAM),实现从图像与领域知识输入到异常定位与掩码生成的自主推理。基于最新多模态IAD基准,该紧凑的30亿参数模型表现超越现有方法,确立新基准。随着MLLM能力持续提升,本研究首次展示基于视觉语言模型的端到端异常检测方案,证实ROAM增强的GRPO具有变革性潜力,为小缺陷数据下的下一代智能检测系统奠定基础。
原文摘要 · Abstract (English)
Industrial Anomaly Detection (IAD) poses a formidable challenge due to the scarcity of defective samples, making it imperative to deploy models capable of robust generalization to detect unseen anomalies effectively. Traditional approaches, often constrained by hand-crafted features or domain-specific expert models, struggle to address this limitation, underscoring the need for a paradigm shift. We introduce AnomalyR1, a pioneering framework that leverages VLM-R1, a Multimodal Large Language Model (MLLM) renowned for its exceptional generalization and interpretability, to revolutionize IAD. By integrating MLLM with Group Relative Policy Optimization (GRPO), enhanced by our novel Reasoned Outcome Alignment Metric (ROAM), AnomalyR1 achieves a fully end-to-end solution that autonomously processes inputs of image and domain knowledge, reasons through analysis, and generates precise anomaly localizations and masks. Based on the latest multimodal IAD benchmark, our compact 3-billion-parameter model outperforms existing methods, establishing state-of-the-art results. As MLLM capabilities continue to advance, this study is the first to deliver an end-to-end VLM-based IAD solution that demonstrates the transformative potential of ROAM-enhanced GRPO, positioning our framework as a forward-looking cornerstone for next-generation intelligent anomaly detection systems in industrial applications with limited defective data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。