arXiv:2603.06313cs.CV2026-03被引 4

用小波分析增强文本提示,提升零样本异常检测能力

WMoE-CLIP: Wavelet-Enhanced Mixture-of-Experts Prompt Learning for Zero-Shot Anomaly Detection

  • 引入小波分解提取多频图像特征,动态优化文本嵌入
  • 通过变分自编码器融合全局语义,增强提示适应性
  • 适合工业与医疗领域中的细微异常检测任务

视觉-语言模型在零样本异常检测(ZSAD)中展现出强大泛化能力,可在无特定任务监督下检测未见异常。然而,现有方法通常依赖固定文本提示,难以捕捉复杂语义,且仅关注空间域特征,限制了对细微异常的检测能力。为此,我们提出一种小波增强的专家混合提示学习方法用于ZSAD。具体而言,采用变分自编码器建模全局语义表示,并将其融入提示以提升对多样异常模式的适应性;小波分解提取多频图像特征,通过跨模态交互动态优化文本嵌入;此外,引入语义感知的专家混合模块聚合上下文信息。在14个工业与医学数据集上的大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Vision-language models have recently shown strong generalization in zero-shot anomaly detection (ZSAD), enabling the detection of unseen anomalies without task-specific supervision. However, existing approaches typically rely on fixed textual prompts, which struggle to capture complex semantics, and focus solely on spatial-domain features, limiting their ability to detect subtle anomalies. To address these challenges, we propose a wavelet-enhanced mixture-of-experts prompt learning method for ZSAD. Specifically, a variational autoencoder is employed to model global semantic representations and integrate them into prompts to enhance adaptability to diverse anomaly patterns. Wavelet decomposition extracts multi-frequency image features that dynamically refine textual embeddings through cross-modal interactions. Furthermore, a semantic-aware mixture-of-experts module is introduced to aggregate contextual information. Extensive experiments on 14 industrial and medical datasets demonstrate the effectiveness of the proposed method.

异常检测小波分析视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。