arXiv:2512.09446cs.CV2025-12

让AI看懂缺陷类型,提升零样本异常检测与分割精度

Defect-aware Hybrid Prompt Optimization via Progressive Tuning for Zero-Shot Multi-type Anomaly Detection and Segmentation

  • 用可学习的提示词融合人工描述与模型自适应特征
  • 图像级检测AUROC提升3.6%,局部定位F1提升5.2%
  • 适合需要识别多种新缺陷的工业质检场景

近期视觉语言模型(如CLIP)在分布偏移下通过文本提示利用高层语义信息,展现出出色的异常检测性能。然而,这些模型常忽略孔洞、切割、划痕等细粒度缺陷线索,且图像与文本之间的模态鸿沟导致细微视觉证据难以被文本准确捕捉。为此,我们通过结构化语义增强“异常”表征,弥合粗粒度异常信号与细粒度缺陷类别之间的差距。提出一种混合提示机制,结合人类可读的缺陷类型描述与可学习的标记嵌入。在此基础上,提出DAPO框架,实现零样本多类型及二值异常检测与分割。DAPO通过学习融合固定文本锚点与可训练标记嵌入的混合缺陷感知提示,对齐异常相关的视觉特征与对应文本语义。在公开基准(MPDD、VisA、MVTec-AD、MAD、Real-IAD)和内部数据集上的实验表明,相比基线模型,DAPO在分布偏移下图像级平均提升3.6%的AUROC与平均精度,在零样本新缺陷定位中平均提升5.2%的AUROC与F1分数。

原文摘要 · Abstract (English)

Recent vision-language models (VLMs) like CLIP have shown impressive anomaly detection performance under significant distribution shift by utilizing high-level semantic information through text prompts. However, these models often overlook fine-grained defect cues, e.g., hole, cut, or scratch, that are essential for understanding the anomaly's nature. Moreover, the modality gap between images and text can lead to subtle visual evidence being poorly captured in textual descriptions. To address the gap, we enhance the representation of "abnormal" with structured semantics, bridging coarse anomaly signals and fine-grained defect categories. We propose a hybrid prompting mechanism that combines human-readable descriptions of defect types with learnable token embeddings. Building on these ideas, we introduce DAPO, a Defect-aware Prompt Optimization framework for zero-shot multi-type and binary anomaly detection and segmentation under distribution shift. DAPO aligns anomaly-relevant visual features with their corresponding textual semantics by learning hybrid defect-aware prompts that combine fixed textual anchors with trainable token embeddings. We conducted experiments on public benchmarks (MPDD, VisA, MVTec-AD, MAD, and Real-IAD) and an internal dataset. The results suggest that compared to the baseline models, DAPO achieves a 3.6% average improvement in AUROC and average precision metrics at the image level under distribution shift, and a 5.2% average improvement in AUROC and F1 when localizing novel anomaly types under zero-shot settings.

异常检测视觉语言模型零样本缺陷识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。