arXiv:2606.29428cs.CV2026-06中稿 · ECCV

在异常数据有限时,仍能精准识别未知缺陷。

Robust Zero-shot Anomaly Detection under Limited Auxiliary Anomaly Priors

论文配图:Robust Zero-shot Anomaly Detection under Limited Auxiliary Anomaly Priors
图 1 · 摘自论文原文
  • 通过文本嵌入注入学习跨域通用异常特征。
  • 在12个数据集上分类准确率提升最高达28.5%。
  • 适合异常样本稀缺的工业质检场景。

零样本异常检测旨在识别任意新领域中的缺陷;然而现有模型假设辅助数据包含丰富的异常多样性,忽视了真实目标领域中更复杂多变的异常模式。本研究提出DIVE,首个针对辅助异常先验有限场景的方法,有效缓解由此导致的性能下降。通过在视觉编码阶段采用浅层与深层文本嵌入注入策略,DIVE学习到辅助训练域与多样化目标域间共享的通用异常概念。同时,提出解耦机制,解决视觉嵌入与对象语义纠缠、且与无关文本提示对齐不佳的问题。实验表明,在辅助数据异常模式有限的设定下,DIVE在两个分类指标上优于当前最优基线最多16.2%和28.5%,在三个分割指标上分别提升23.4%、24.1%和47.0%,十二个数据集平均表现显著领先。此外,当辅助数据具备充分异常多样性时,其性能仍保持高度竞争力。

原文摘要 · Abstract (English)

Zero-shot anomaly detection aims to identify defects in arbitrary novel domains; however, existing models assume that the auxiliary data contains a rich diversity of anomalies, neglecting the far more complex and unpredictable variations in real-world target domains. This study introduces DIVE, the first approach to investigate the scenario of limited auxiliary anomaly priors and resolve the resulting substantial performance degradation. Through a shallow-and-deep text embedding injection strategy during visual encoding, DIVE learns to abstract generic anomaly concepts shared across the auxiliary training domain and diverse target domains. Moreover, we propose a disentanglement mechanism to tackle the suboptimal alignment between visual embeddings entangled with object semantics and object-agnostic textual prompts. Experiments demonstrate that, under the setting of limited anomaly patterns in auxiliary data, DIVE outperforms SOTA baselines by up to 16.2% and 28.5% on two classification metrics, and 23.4%, 24.1%, and 47.0% on three segmentation metrics, in terms of average performance across twelve datasets. Furthermore, it maintains highly competitive performance when auxiliary data exhibits sufficient anomaly diversity.

异常检测零样本工业质检文本注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。