arXiv:2503.12910cs.CV2025-03被引 2

用图像信息修正文本描述,让模型更准识别工业和医疗中的异常

AFR-CLIP: Enhancing Zero-Shot Industrial Anomaly Detection with Stateless-to-Stateful Anomaly Feature Rectification

  • 用图像引导文本重构,把缺陷信息注入无状态提示词
  • 在多个数据集上实现当前最优零样本异常检测性能
  • 适合需要不依赖标注数据的工业质检与医学影像分析场景

近期,零样本异常检测(ZSAD)已成为工业检测和医学诊断的关键范式,可在无需目标数据集样本的情况下检测新物体的缺陷。现有基于CLIP的ZSAD方法通过测量视觉与文本特征之间的余弦相似度生成异常图。然而,CLIP对物体类别而非异常状态的对齐限制了其在异常检测中的有效性。为此,我们提出AFR-CLIP,一种基于CLIP的异常特征修正框架。AFR-CLIP首先进行图像引导的文本修正,将图像中隐含的缺陷信息嵌入到仅描述物体类别的无状态提示词中。随后,将增强后的文本嵌入与两个预定义的状态嵌入(正常或异常)进行比较,其文本间相似度生成突出缺陷区域的异常图。为进一步提升对多尺度特征和复杂异常的感知能力,引入自提示(SP)和多块特征聚合(MPFA)模块。在涵盖工业与医学领域的十一个异常检测基准上进行了大量实验,验证了AFR-CLIP在零样本异常检测中的优越性。

原文摘要 · Abstract (English)

Recently, zero-shot anomaly detection (ZSAD) has emerged as a pivotal paradigm for industrial inspection and medical diagnostics, detecting defects in novel objects without requiring any target-dataset samples during training. Existing CLIP-based ZSAD methods generate anomaly maps by measuring the cosine similarity between visual and textual features. However, CLIP's alignment with object categories instead of their anomalous states limits its effectiveness for anomaly detection. To address this limitation, we propose AFR-CLIP, a CLIP-based anomaly feature rectification framework. AFR-CLIP first performs image-guided textual rectification, embedding the implicit defect information from the image into a stateless prompt that describes the object category without indicating any anomalous state. The enriched textual embeddings are then compared with two pre-defined stateful (normal or abnormal) embeddings, and their text-on-text similarity yields the anomaly map that highlights defective regions. To further enhance perception to multi-scale features and complex anomalies, we introduce self prompting (SP) and multi-patch feature aggregation (MPFA) modules. Extensive experiments are conducted on eleven anomaly detection benchmarks across industrial and medical domains, demonstrating AFR-CLIP's superiority in ZSAD.

异常检测CLIP零样本工业质检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。