arXiv:2508.16157cs.CVcs.AI2025-08

无需人工提示,自动学习异常检测的可调提示。

Beyond Human-prompting: Adaptive Prompt Tuning with Semantic Alignment for Anomaly Detection

  • 用噪声生成伪异常样本,训练可学习提示以捕捉场景相关异常。
  • 在MVTec AD等数据集上达到最新最好性能,少样本下表现优异。
  • 适合无标注异常数据、需快速适配新场景的应用场景。

预训练视觉语言模型在异常检测中展现出潜力,但现有方法受限于人工设计提示和缺乏可访问的异常样本,导致上下文相关的异常理解不足。本文提出自适应提示调优框架APT,无需先验知识,支持少样本学习。APT通过噪声扰动生成自生异常样本,训练可学习提示以捕获不同场景下的上下文依赖异常。为防止对合成噪声过拟合,引入自优化元提示引导机制(SMGS),迭代对齐提示与通用异常语义,并融合多样化合成异常。该系统不仅提升像素级异常检测效果,还在多个基准数据集(如MVTec AD)上实现领先性能,无需人工提示设计,为真实场景下的异常检测提供鲁棒且通用的解决方案。

原文摘要 · Abstract (English)

Pre-trained Vision-Language Models (VLMs) have recently shown promise in detecting anomalies. However, previous approaches are fundamentally limited by their reliance on human-designed prompts and the lack of accessible anomaly samples, leading to significant gaps in context-specific anomaly understanding. In this paper, we propose \textbf{A}daptive \textbf{P}rompt \textbf{T}uning with semantic alignment for anomaly detection (APT), a groundbreaking prior knowledge-free, few-shot framework and overcomes the limitations of traditional prompt-based approaches. APT uses self-generated anomaly samples with noise perturbations to train learnable prompts that capture context-dependent anomalies in different scenarios. To prevent overfitting to synthetic noise, we propose a Self-Optimizing Meta-prompt Guiding Scheme (SMGS) that iteratively aligns the prompts with general anomaly semantics while incorporating diverse synthetic anomaly. Our system not only advances pixel-wise anomaly detection, but also achieves state-of-the-art performance on multiple benchmark datasets without requiring prior knowledge for prompt crafting, establishing a robust and versatile solution for real-world anomaly detection.

异常检测提示调优少样本学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。