arXiv:2607.29370cs.CV2026-07

通过语义提示与频域自适应学习,提升零样本异常检测的精度与定位能力

VFAD: Variational Semantic Prompting Meets Frequency-Adaptive Representation Learning for Zero-Shot Anomaly Detection

论文配图:VFAD: Variational Semantic Prompting Meets Frequency-Adaptive Representation Learning for Zero-Shot Anomaly Detection
图 1 · 摘自论文原文
  • 引入变分语义提示提取器,动态聚合局部视觉语义并增强跨模态对齐
  • 设计频域自适应表征聚合模块,利用小波分解捕捉异常特异性特征
  • 在13个工业与医疗数据集上超越现有方法,适合需要高精度定位的场景

零样本异常检测(ZSAD)旨在无需目标类别训练数据的情况下检测和定位未见类别的异常。尽管基于CLIP的方法通过视觉-语言对齐展现出良好泛化能力,但仍难以捕捉多样的异常语义和细微的局部变化。为此,我们提出VFAD,一个融合变分语义提示与频域自适应表征学习的统一框架。具体地,提出变分语义提示提取器(VSPE),从密集补丁标记中自适应聚合相关异常的局部语义,并通过变分信息瓶颈进行正则化,从而融入细粒度视觉线索,实现更精确的跨模态对齐。此外,设计频域自适应表征聚合(FARA)模块,利用小波基频率分解与频段专属专家聚合,增强异常判别性视觉表征。通过联合强化语义引导与视觉表示学习,VFAD显著提升异常判别力与细粒度定位能力。在13个工业与医疗基准上的大量实验表明,该方法在多种异常场景下持续优于现有最先进方法。代码将在发表后公开。

原文摘要 · Abstract (English)

Zero-shot anomaly detection (ZSAD) aims to detect and localize anomalies in unseen categories without access to target-specific training data. Although recent CLIP-based methods have demonstrated promising generalization through vision-language alignment, they remain limited in capturing diverse anomaly semantics and subtle local variations. To address these limitations, we propose VFAD, a unified framework that combines variational semantic prompting with frequency-adaptive representation learning. Specifically, we introduce a Variational Semantic Prompt Extractor (VSPE), which adaptively aggregates anomaly-relevant local semantics from dense patch tokens and regularizes them through a variational information bottleneck, thereby incorporating fine-grained visual cues and enabling more precise cross-modal alignment. Furthermore, we develop a Frequency-Adaptive Representation Aggregation (FARA) module that leverages wavelet-based frequency decomposition and frequency-specific expert aggregation to enhance anomaly-discriminative visual representations. By jointly strengthening semantic guidance and visual representation learning, VFAD improves both anomaly discrimination and fine-grained localization. Extensive experiments on 13 industrial and medical benchmarks demonstrate that VFAD consistently outperforms existing state-of-the-art ZSAD methods across diverse anomaly scenarios. The code will be publicly available upon publication.

异常检测零样本视觉语言模型频域学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。