用离散潜空间掩码建模检测图像逻辑异常,突破局部模式限制。
LADMIM: Logical Anomaly Detection with Masked Image Modeling in Discrete Latent Space
- 将异常检测转为掩码区域潜变量重建任务,捕捉全局依赖关系。
- 在五个基准上表现优异,无需预训练分割模型即可达成可比性能。
- 适合关注图像逻辑一致性与无监督异常检测的研究者。
检测物体组合错误或位置偏移等逻辑异常是无监督异常检测(AD)中的难题。传统方法主要关注正常图像的局部模式,难以发现全局模式中的逻辑异常。为此,我们提出逻辑异常检测的掩码图像建模框架(LADMIM),利用掩码图像建模与离散表示学习的结合。核心思想是:通过预测缺失区域的离散潜变量分布,迫使模型学习图像块间的长距离依赖关系。由于离散潜变量分布对像素空间的低层差异具有不变性,模型能聚焦于图像的逻辑依赖结构,从而提升逻辑异常检测精度。我们在五个基准数据集上评估了性能,结果表明该方法在不依赖任何预训练分割模型的前提下实现了可比的性能。此外,通过全面实验揭示了影响逻辑异常检测的关键因素。
原文摘要 · Abstract (English)
Detecting anomalies such as an incorrect combination of objects or deviations in their positions is a challenging problem in unsupervised anomaly detection (AD). Since conventional AD methods mainly focus on local patterns of normal images, they struggle with detecting logical anomalies that appear in the global patterns. To effectively detect these challenging logical anomalies, we introduce Logical Anomaly Detection with Masked Image Modeling (LADMIM), a novel unsupervised AD framework that harnesses the power of masked image modeling and discrete representation learning. Our core insight is that predicting the missing region forces the model to learn the long-range dependencies between patches. Specifically, we formulate AD as a mask completion task, which predicts the distribution of discrete latents in the masked region. As a distribution of discrete latents is invariant to the low-level variance in the pixel space, the model can desirably focus on the logical dependencies in the image, which improves accuracy in the logical AD. We evaluate the AD performance on five benchmarks and show that our approach achieves compatible performance without any pre-trained segmentation models. We also conduct comprehensive experiments to reveal the key factors that influence logical AD performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。