arXiv:2601.09147cs.CVcs.AI2026-01

通过融合多尺度视觉特征,提升工业缺陷检测的细粒度感知能力。

SSVP: Synergistic Semantic-Visual Prompting for Industrial Zero-Shot Anomaly Detection

  • 设计双机制融合多视觉编码,增强模型对结构细节的敏感度。
  • 在MVTec-AD数据集上达到93.0%图像级与92.2%像素级AUROC。
  • 适合追求高精度零样本工业质检的工程师与研究者。

零样本异常检测(ZSAD)利用视觉语言模型实现无需标注的工业质检。现有方法受限于单一视觉主干网络,难以兼顾全局语义泛化与细粒度结构判别。为此,我们提出协同语义-视觉提示(SSVP),高效融合多种视觉编码以提升模型细粒度感知能力。具体而言,SSVP引入分层语义-视觉协同(HSVS)机制,将DINOv3的多尺度结构先验深度融入CLIP语义空间;随后,视觉条件提示生成器(VCPG)通过跨模态注意力动态生成提示,使语言查询精准锚定异常模式。此外,为缓解全局评分与局部证据之间的偏差,视觉-文本异常映射器(VTAM)构建双门控校准范式。在七个工业基准上的广泛评估验证了方法的鲁棒性:在MVTec-AD上,SSVP实现93.0%图像级AUROC与92.2%像素级AUROC,显著优于现有零样本方法。

原文摘要 · Abstract (English)

Zero-Shot Anomaly Detection (ZSAD) leverages Vision-Language Models (VLMs) to enable supervision-free industrial inspection. However, existing ZSAD paradigms are constrained by single visual backbones, which struggle to balance global semantic generalization with fine-grained structural discriminability. To bridge this gap, we propose Synergistic Semantic-Visual Prompting (SSVP), that efficiently fuses diverse visual encodings to elevate model's fine-grained perception. Specifically, SSVP introduces the Hierarchical Semantic-Visual Synergy (HSVS) mechanism, which deeply integrates DINOv3's multi-scale structural priors into the CLIP semantic space. Subsequently, the Vision-Conditioned Prompt Generator (VCPG) employs cross-modal attention to guide dynamic prompt generation, enabling linguistic queries to precisely anchor to specific anomaly patterns. Furthermore, to address the discrepancy between global scoring and local evidence, the Visual-Text Anomaly Mapper (VTAM) establishes a dual-gated calibration paradigm. Extensive evaluations on seven industrial benchmarks validate the robustness of our method; SSVP achieves state-of-the-art performance with 93.0% Image-AUROC and 92.2% Pixel-AUROC on MVTec-AD, significantly outperforming existing zero-shot approaches.

异常检测零样本视觉语言模型工业质检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。