用视觉语言模型实现可解释的异常检测,效果超越现有方法
LogicAD: Explainable Anomaly Detection via VLM-based Text Feature Extraction
- 结合视觉语言模型与逻辑推理机制,自动提取图像语义特征
- 在MVTec LOCO数据集上达到86.0%的AUROC和83.7%的F1-max
- 不仅能准确识别异常,还能提供人类可理解的解释,适合工业质检场景
逻辑图像理解涉及对图像内容间关系与一致性的解读与推理,对工业质检等应用至关重要。传统异常检测依赖先验知识设计算法,需大量人工标注、计算资源和训练数据。自回归多模态视觉语言模型(AVLMs)在跨领域视觉推理中表现优异,但尚未应用于逻辑异常检测。本文探索利用AVLMs进行逻辑异常检测,结合格式嵌入与逻辑推理器,在公开基准MVTec LOCO AD上实现86.0%的AUROC与83.7%的F1-max,性能显著优于现有最先进方法,并能生成异常解释。
原文摘要 · Abstract (English)
Logical image understanding involves interpreting and reasoning about the relationships and consistency within an image's visual content. This capability is essential in applications such as industrial inspection, where logical anomaly detection is critical for maintaining high-quality standards and minimizing costly recalls. Previous research in anomaly detection (AD) has relied on prior knowledge for designing algorithms, which often requires extensive manual annotations, significant computing power, and large amounts of data for training. Autoregressive, multimodal Vision Language Models (AVLMs) offer a promising alternative due to their exceptional performance in visual reasoning across various domains. Despite this, their application to logical AD remains unexplored. In this work, we investigate using AVLMs for logical AD and demonstrate that they are well-suited to the task. Combining AVLMs with format embedding and a logic reasoner, we achieve SOTA performance on public benchmarks, MVTec LOCO AD, with an AUROC of 86.0% and F1-max of 83.7%, along with explanations of anomalies. This significantly outperforms the existing SOTA method by a large margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。