arXiv:2412.00890cs.CV2024-12被引 10

用视觉语言模型提升工业缺陷检测的准确率与可解释性

Exploring Large Vision-Language Models for Robust and Efficient Industrial Anomaly Detection

  • 通过对比学习对齐图文特征,让正常样本聚类、异常分离
  • 在MVTec-AD和VisA数据集上超越现有最佳方法
  • 能精确定位异常区域,适合工业质检场景

工业异常检测在制造过程维护与质量控制中至关重要。本文提出一种新方法——基于对比跨模态训练的视觉语言异常检测(CLAD),利用大视觉语言模型(LVLMs)提升工业场景下的异常检测与定位能力。CLAD通过对比学习将视觉与文本特征映射到统一嵌入空间,使正常样本聚集而异常样本分离。在两个基准工业数据集MVTec-AD和VisA上的大量实验表明,CLAD在图像级异常检测与像素级异常定位任务中均优于当前最优方法。此外,通过消融实验与人工评估验证了关键组件的有效性。该方法不仅性能优异,还通过精准定位增强了可解释性,为实际工业应用提供了有前景的解决方案。

原文摘要 · Abstract (English)

Industrial anomaly detection (IAD) plays a crucial role in the maintenance and quality control of manufacturing processes. In this paper, we propose a novel approach, Vision-Language Anomaly Detection via Contrastive Cross-Modal Training (CLAD), which leverages large vision-language models (LVLMs) to improve both anomaly detection and localization in industrial settings. CLAD aligns visual and textual features into a shared embedding space using contrastive learning, ensuring that normal instances are grouped together while anomalies are pushed apart. Through extensive experiments on two benchmark industrial datasets, MVTec-AD and VisA, we demonstrate that CLAD outperforms state-of-the-art methods in both image-level anomaly detection and pixel-level anomaly localization. Additionally, we provide ablation studies and human evaluation to validate the importance of key components in our method. Our approach not only achieves superior performance but also enhances interpretability by accurately localizing anomalies, making it a promising solution for real-world industrial applications.

异常检测视觉语言模型工业质检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。