arXiv:2507.22692cs.CV2025-07ICCV被引 4

用扩散模型的去噪轨迹检测图像异常,无需再训练。

Zero-Shot Image Anomaly Detection Using Generative Foundation Models

  • 用扩散模型的去噪过程提取纹理与语义信息作为感知模板。
  • 在多个数据集上达到接近完美的异常检测性能,部分任务仍有提升空间。
  • 仅需在CelebA上训练一次,即可跨数据集应用,适合部署于开放世界场景。

在开放世界环境中部署安全视觉系统的关键在于检测分布外(OOD)输入。本文重新审视扩散模型,不将其作为生成器,而是作为通用感知模板用于OOD检测。研究探索了基于得分的生成模型作为跨未见数据集进行语义异常检测的基础工具。具体而言,我们利用去噪扩散模型(DDMs)的去噪轨迹作为丰富的纹理与语义信息来源。通过分析经结构相似性指数(SSIM)放大的Stein得分误差,提出一种无需在目标数据集上再训练的新型异常检测方法。该方法在多个基准上超越现有最佳结果,且仅需在单一数据集——CelebA上进行训练,其表现甚至优于更常用的数据集ImageNet。实验结果显示,在某些基准上接近完美性能,其他任务仍具提升空间,凸显生成式基础模型在异常检测中的潜力。

原文摘要 · Abstract (English)

Detecting out-of-distribution (OOD) inputs is pivotal for deploying safe vision systems in open-world environments. We revisit diffusion models, not as generators, but as universal perceptual templates for OOD detection. This research explores the use of score-based generative models as foundational tools for semantic anomaly detection across unseen datasets. Specifically, we leverage the denoising trajectories of Denoising Diffusion Models (DDMs) as a rich source of texture and semantic information. By analyzing Stein score errors, amplified through the Structural Similarity Index Metric (SSIM), we introduce a novel method for identifying anomalous samples without requiring re-training on each target dataset. Our approach improves over state-of-the-art and relies on training a single model on one dataset -- CelebA -- which we find to be an effective base distribution, even outperforming more commonly used datasets like ImageNet in several settings. Experimental results show near-perfect performance on some benchmarks, with notable headroom on others, highlighting both the strength and future potential of generative foundation models in anomaly detection.

异常检测扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。