用多模态大模型做零样本异常检测,提升对图像异常的精准识别与推理能力。
Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models
- 设计双看特征匹配机制,自适应聚焦异常视觉信息。
- 在12.5万条指令数据上训练,显著优于通用模型。
- 适合医疗、3D等数据稀缺场景的异常检测研究者使用。
零样本异常检测(ZSAD)是一种新兴范式,无需大量正常样本即可应对真实世界中数据受限的场景。尽管多模态大语言模型(MLLMs)在视觉任务中展现出强大推理能力,但其对图像异常的判断仍缺乏系统研究。为此,我们构建了首个视觉指令微调数据集Anomaly-Instruct-125k和评估基准VisA-D&R。实验发现,GPT-4o等现有MLLM难以准确识别图像中的细粒度异常。为此,提出Anomaly-OneVision(Anomaly-OV),首个专注于零样本异常检测与推理的视觉助手。受人类视觉检查启发,Anomaly-OV采用双看特征匹配(LTFM)机制,动态选择并强化异常视觉标记。大量实验表明,Anomaly-OV在检测与推理性能上显著超越先进通用模型。研究还拓展至医学和3D异常检测,为未来工作提供方向。
原文摘要 · Abstract (English)
Zero-Shot Anomaly Detection (ZSAD) is an emerging AD paradigm. Unlike the traditional unsupervised AD setting that requires a large number of normal samples to train a model, ZSAD is more practical for handling data-restricted real-world scenarios. Recently, Multimodal Large Language Models (MLLMs) have shown revolutionary reasoning capabilities in various vision tasks. However, the reasoning of image abnormalities remains underexplored due to the lack of corresponding datasets and benchmarks. To facilitate research in AD & reasoning, we establish the first visual instruction tuning dataset, Anomaly-Instruct-125k, and the evaluation benchmark, VisA-D&R. Through investigation with our benchmark, we reveal that current MLLMs like GPT-4o cannot accurately detect and describe fine-grained anomalous details in images. To address this, we propose Anomaly-OneVision (Anomaly-OV), the first specialist visual assistant for ZSAD and reasoning. Inspired by human behavior in visual inspection, Anomaly-OV leverages a Look-Twice Feature Matching (LTFM) mechanism to adaptively select and emphasize abnormal visual tokens. Extensive experiments demonstrate that Anomaly-OV achieves significant improvements over advanced generalist models in both detection and reasoning. Extensions to medical and 3D AD are provided for future study. The link to our project page: https://xujiacong.github.io/Anomaly-OV/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。