通过视觉校准语义距离,提升零样本异常检测的准确性
Proximity-CLIP: Text-Guided Semantic Proximity Learning for Zero-Shot Anomaly Detection

- 用动态正则化学习合适的语义间距,兼顾区分度与结构一致性
- 引入异常查询模块,精准定位局部缺陷特征,避免微小异常被掩盖
- 仅需少量修改架构,即可在多个基准上超越现有最佳方法
视觉-语言模型为零样本异常检测(ZSAD)提供了前景广阔的解决方案。然而,由于以物体为中心的偏见,正常与异常文本原型存在高度语义重叠。尽管强制两者正交可增强区分性,但将高度连续的视觉输入映射到截然不同的原型会引发几何困境,破坏预训练的结构连续性。为此,我们提出Proximity-CLIP框架,通过视觉校准语义边界引导视觉适应。首先,设计一种视觉校准的语义邻近学习机制,利用有界动态正则化学习合理语义间距,确保判别分离的同时保留结构对齐。其次,构建基于文本先验的异常查询模块(AQM),以校准后的异常原型作为语义查询,主动从上下文视觉块中检索局部缺陷线索,缓解全局池化过程中细微异常的稀释问题。大量实验表明,Proximity-CLIP在多个ZSAD基准上均优于当前最优方法,且仅需极少架构改动。
原文摘要 · Abstract (English)
Vision-language models offer a promising approach for zero-shot anomaly detection (ZSAD). However, due to object-centric bias, normal and anomalous text prototypes exhibit a high semantic overlap. While enforcing strict orthogonality between them improves discriminability, mapping highly contiguous visual inputs onto drastically orthogonal prototypes introduces a geometric dilemma, disrupting the pre-trained structural continuity. To address this problem, we propose Proximity-CLIP, a framework that visually calibrates the semantic margin to guide visual adaptation. First, we introduce a visually-calibrated semantic proximity learning mechanism that uses a bounded dynamic regularization to learn an appropriate semantic margin, ensuring discriminative separation while preserving structural alignment. Second, we design an Anomaly Query Module (AQM) driven by these text priors. Using the calibrated anomalous prototype as a semantic query, the AQM actively retrieves localized defect cues from contextual visual patches, mitigating the dilution of subtle anomalies during global pooling. Extensive experiments demonstrate that Proximity-CLIP outperforms current state-of-the-art methods across multiple ZSAD benchmarks with minimal architectural modifications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。