arXiv:2505.17692cs.CVcs.AI2025-05被引 2

用视觉感知生成提示,让模型在无训练数据下精准识别异常

ViP$^2$-CLIP: Visual-Perception Prompting with Unified Alignment for Zero-Shot Anomaly Detection

  • 根据图像全局与多尺度局部特征自动生成细粒度提示
  • 在15个工业和医疗数据集上达到当前最佳性能
  • 无需类别标签,适合标签模糊或隐私受限场景

零样本异常检测(ZSAD)旨在不依赖目标域训练样本的情况下检测异常,仅依靠外部辅助数据。现有基于CLIP的方法通过手工设计或静态可学习提示激活模型的ZSAD潜力,前者工程成本高且语义覆盖有限,后者对所有异常类型使用相同描述,无法适应复杂变化。此外,由于CLIP最初在大规模分类任务上预训练,其异常分割质量高度依赖类别名称的精确措辞,严重限制了依赖类别标签的提示策略。为此,我们提出ViP²-CLIP。核心思想是视觉感知提示(ViP-Prompt)机制,融合全局与多尺度局部视觉上下文,自适应生成细粒度文本提示,消除人工模板和类别名称先验。该设计使模型能聚焦于精确的异常区域,尤其适用于类别标签模糊或隐私受限场景。在15个工业和医疗基准上的大量实验表明,ViP²-CLIP实现了最先进的性能和强跨域泛化能力。

原文摘要 · Abstract (English)

Zero-shot anomaly detection (ZSAD) aims to detect anomalies without any target domain training samples, relying solely on external auxiliary data. Existing CLIP-based methods attempt to activate the model's ZSAD potential via handcrafted or static learnable prompts. The former incur high engineering costs and limited semantic coverage, whereas the latter apply identical descriptions across diverse anomaly types, thus fail to adapt to complex variations. Furthermore, since CLIP is originally pretrained on large-scale classification tasks, its anomaly segmentation quality is highly sensitive to the exact wording of class names, severely constraining prompting strategies that depend on class labels. To address these challenges, we introduce ViP$^{2}$-CLIP. The key insight of ViP$^{2}$-CLIP is a Visual-Perception Prompting (ViP-Prompt) mechanism, which fuses global and multi-scale local visual context to adaptively generate fine-grained textual prompts, eliminating manual templates and class-name priors. This design enables our model to focus on precise abnormal regions, making it particularly valuable when category labels are ambiguous or privacy-constrained. Extensive experiments on 15 industrial and medical benchmarks demonstrate that ViP$^{2}$-CLIP achieves state-of-the-art performance and robust cross-domain generalization.

异常检测CLIP提示学习零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。