用视觉增强的多模态大模型实现零样本缺陷检测,提升细粒度识别能力。
VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection
- 引入缺陷敏感结构学习,利用视觉分支传递局部相似性信息增强异常判别
- 设计局域性标记压缩投影器,挖掘多层级局部特征提升细粒度定位精度
- 构建真实工业异常检测数据集RIAD,支持大模型在开放场景下的异常分析
零样本异常检测(ZSAD)通过建立文本提示与检测图像之间的特征映射,识别未见过物体中的异常,在柔性工业制造中具有重要研究价值。然而,现有方法受限于封闭世界设定,难以应对预定义提示外的未知缺陷。近期,将多模态大语言模型(MLLM)应用于工业异常检测(IAD)成为可行方案。与固定提示方法不同,MLLM具备开放式文本理解能力,可实现更灵活的异常分析。但其面临挑战:异常常表现为细微区域变化,与正常样本视觉差异极小。为此,本文提出新型框架VMAD(视觉增强的MLLM异常检测),通过注入视觉引导的IAD知识与细粒度感知能力,实现精准检测与全面分析。具体地,设计缺陷敏感结构学习机制,将视觉分支的块相似性线索迁移至MLLM以增强异常区分;提出新颖的局域性标记压缩投影器,从局部上下文中挖掘多层次特征以提升细粒度检测;此外,构建真实工业异常检测数据集RIAD,包含详尽异常描述与分析,为基于MLLM的IAD发展提供宝贵资源。在多个零样本基准测试(包括MVTec-AD、Visa、WFDD和RIAD)上,实验结果表明,本方法显著优于现有最先进方法。
原文摘要 · Abstract (English)
Zero-shot anomaly detection (ZSAD) recognizes and localizes anomalies in previously unseen objects by establishing feature mapping between textual prompts and inspection images, demonstrating excellent research value in flexible industrial manufacturing. However, existing ZSAD methods are limited by closed-world settings, struggling to unseen defects with predefined prompts. Recently, adapting Multimodal Large Language Models (MLLMs) for Industrial Anomaly Detection (IAD) presents a viable solution. Unlike fixed-prompt methods, MLLMs exhibit a generative paradigm with open-ended text interpretation, enabling more adaptive anomaly analysis. However, this adaption faces inherent challenges as anomalies often manifest in fine-grained regions and exhibit minimal visual discrepancies from normal samples. To address these challenges, we propose a novel framework VMAD (Visual-enhanced MLLM Anomaly Detection) that enhances MLLM with visual-based IAD knowledge and fine-grained perception, simultaneously providing precise detection and comprehensive analysis of anomalies. Specifically, we design a Defect-Sensitive Structure Learning scheme that transfers patch-similarities cues from visual branch to our MLLM for improved anomaly discrimination. Besides, we introduce a novel visual projector, Locality-enhanced Token Compression, which mines multi-level features in local contexts to enhance fine-grained detection. Furthermore, we introduce the Real Industrial Anomaly Detection (RIAD), a comprehensive IAD dataset with detailed anomaly descriptions and analyses, offering a valuable resource for MLLM-based IAD development. Extensive experiments on zero-shot benchmarks, including MVTec-AD, Visa, WFDD, and RIAD datasets, demonstrate our superior performance over state-of-the-art methods. The code and dataset will be available soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。