arXiv:2606.09670cs.CVcs.AI2026-06

用视觉提示与双教师机制提升异常检测在真实场景下的泛化能力

Visual Prompting Meets Feature Reconstruction-Based Anomaly Detection with Dual-Teacher Supervision

论文配图:Visual Prompting Meets Feature Reconstruction-Based Anomaly Detection with Dual-Teacher Supervision
图 1 · 摘自论文原文
  • 通过前景背景掩码实现物体隔离,增强对复杂场景的适应性
  • 解冻教师模型并结合扩散生成数据,使检测性能提升3.5个百分点
  • 适合需要高鲁棒性的工业缺陷检测场景

近期异常检测方法在MVTec等标准数据集上取得了完美检测与分割效果,但当物体尺度、视角、背景、光照或位置一致性等基础假设被打破时,性能显著下降,难以应用于真实场景。为此,本文提出三项关键贡献:(1) 基于前景-背景掩码的视觉提示流程,有效分离目标物体;(2) 解冻学生-教师模型中的教师网络以提升领域适应性;(3) 利用扩散模型生成合成图像进行数据增强。采用以Masked Multiscale Reconstruction(MMR)为骨干网络的方法,在挑战性强的AeBAD数据集上相较此前最优方法提升3.5个百分点。

原文摘要 · Abstract (English)

Recent Anomaly Detection methods achieve perfect detection and segmentation scores on well-established datasets, such as MVTec. However, many of these methods face challenges when foundational assumptions - such as consistent object scale, viewpoint, background, illumination, and centered placement - are violated. Those variations that occur render anomaly detection methods unusable in many real-world scenarios. To address these limitations, we introduce three key contributions: (1) a visual prompting pipeline that isolates objects using foreground-background masking; (2) a mechanism for unfreezing the teacher in student-teacher models to improve domain adaptability; and (3) a data augmentation strategy leveraging diffusion-generated synthetic images to enhance anomaly detection performance. We achieve a 3.5 percentage point improvement over the previous state-of-the-art on the challenging AeBAD dataset by using the Masked Multiscale Reconstruction (MMR) model as our backbone.

异常检测视觉提示扩散模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。