arXiv:2511.18116cs.CV2025-11AAAI被引 6

用视觉引导的提示组合,实现跨领域异常检测新突破

PromptMoE: Generalizable Zero-Shot Anomaly Detection via Visually-Guided Prompt Mixtures

  • 构建专家提示池与视觉门控的混合专家机制
  • 在15个工业医疗数据集上达顶尖性能
  • 适合需要零样本异常检测的工程与医学场景

零样本异常检测(ZSAD)旨在识别未见物体类别的图像中异常区域。尽管基于视觉语言模型如CLIP的方法展现潜力,但其性能受限于现有提示工程策略。当前方法依赖单一固定、可学习或密集动态提示,存在表征瓶颈且易在辅助数据上过拟合,难以泛化到未见异常的复杂多样性。为此,我们提出PromptMoE。核心思想是采用组合式提示学习:不学习单一提示,而是学习一组专家提示作为可组合的语义基元,并通过视觉引导的混合专家(MoE)机制动态组合它们。框架通过视觉引导的提示混合(VGMoP)实现,利用图像门控的稀疏MoE聚合多样正常与异常专家状态提示,生成语义丰富的文本表示并具备强泛化能力。在15个工业与医疗领域的数据集上进行的大量实验验证了PromptMoE的有效性与先进性能。

原文摘要 · Abstract (English)

Zero-Shot Anomaly Detection (ZSAD) aims to identify and localize anomalous regions in images of unseen object classes. While recent methods based on vision-language models like CLIP show promise, their performance is constrained by existing prompt engineering strategies. Current approaches, whether relying on single fixed, learnable, or dense dynamic prompts, suffer from a representational bottleneck and are prone to overfitting on auxiliary data, failing to generalize to the complexity and diversity of unseen anomalies. To overcome these limitations, we propose $\mathtt{PromptMoE}$. Our core insight is that robust ZSAD requires a compositional approach to prompt learning. Instead of learning monolithic prompts, $\mathtt{PromptMoE}$ learns a pool of expert prompts, which serve as a basis set of composable semantic primitives, and a visually-guided Mixture-of-Experts (MoE) mechanism to dynamically combine them for each instance. Our framework materializes this concept through a Visually-Guided Mixture of Prompt (VGMoP) that employs an image-gated sparse MoE to aggregate diverse normal and abnormal expert state prompts, generating semantically rich textual representations with strong generalization. Extensive experiments across 15 datasets in industrial and medical domains demonstrate the effectiveness and state-of-the-art performance of $\mathtt{PromptMoE}$.

异常检测零样本提示学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。