通过语义提示与频域自适应学习,提升零样本异常检测的精度与定位能力
VFAD: Variational Semantic Prompting Meets Frequency-Adaptive Representation Learning for Zero-Shot Anomaly Detection

- 引入变分语义提示提取器,动态聚合局部视觉语义并增强跨模态对齐
- 设计频域自适应表征聚合模块,利用小波分解捕捉异常特异性特征
- 在13个工业与医疗数据集上超越现有方法,适合需要高精度定位的场景
零样本异常检测(ZSAD)旨在无需目标类别训练数据的情况下检测和定位未见类别的异常。尽管基于CLIP的方法通过视觉-语言对齐展现出良好泛化能力,但仍难以捕捉多样的异常语义和细微的局部变化。为此,我们提出VFAD,一个融合变分语义提示与频域自适应表征学习的统一框架。具体地,提出变分语义提示提取器(VSPE),从密集补丁标记中自适应聚合相关异常的局部语义,并通过变分信息瓶颈进行正则化,从而融入细粒度视觉线索,实现更精确的跨模态对齐。此外,设计频域自适应表征聚合(FARA)模块,利用小波基频率分解与频段专属专家聚合,增强异常判别性视觉表征。通过联合强化语义引导与视觉表示学习,VFAD显著提升异常判别力与细粒度定位能力。在13个工业与医疗基准上的大量实验表明,该方法在多种异常场景下持续优于现有最先进方法。代码将在发表后公开。
原文摘要 · Abstract (English)
Zero-shot anomaly detection (ZSAD) aims to detect and localize anomalies in unseen categories without access to target-specific training data. Although recent CLIP-based methods have demonstrated promising generalization through vision-language alignment, they remain limited in capturing diverse anomaly semantics and subtle local variations. To address these limitations, we propose VFAD, a unified framework that combines variational semantic prompting with frequency-adaptive representation learning. Specifically, we introduce a Variational Semantic Prompt Extractor (VSPE), which adaptively aggregates anomaly-relevant local semantics from dense patch tokens and regularizes them through a variational information bottleneck, thereby incorporating fine-grained visual cues and enabling more precise cross-modal alignment. Furthermore, we develop a Frequency-Adaptive Representation Aggregation (FARA) module that leverages wavelet-based frequency decomposition and frequency-specific expert aggregation to enhance anomaly-discriminative visual representations. By jointly strengthening semantic guidance and visual representation learning, VFAD improves both anomaly discrimination and fine-grained localization. Extensive experiments on 13 industrial and medical benchmarks demonstrate that VFAD consistently outperforms existing state-of-the-art ZSAD methods across diverse anomaly scenarios. The code will be publicly available upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。