arXiv:2503.06201cs.CVcs.AI2025-03AAAI被引 11

通过扩散模型中间步骤检测生成图像,准确率超95%

Explainable Synthetic Image Detection through Diffusion Timestep Ensembling

  • 利用扩散模型反演过程中的多步噪声图像特征进行检测
  • 在常规和困难样本上分别达98.91%与95.89%准确率
  • 提供可解释的假图缺陷分析,适合安全检测场景

扩散模型的进展使得生成图像愈发逼真,带来严重安全风险。本研究实证发现,DDIM反演过程中不同时间步会暴露合成图像与真实图像间的细微差异,如傅里叶高频能量分布和像素间方差特征。基于此,提出一种新检测方法:直接利用多步噪声图像特征训练集成模型,跳过传统重建策略。为提升可解释性,引入基于指标的解释生成与优化模块,识别并阐释生成缺陷。此外,构建了GenHard和GenExplain两个基准数据集,包含更难检测的样本及高质量理由。大量实验表明,该方法在常规与挑战样本上分别达到98.91%和95.89%检测准确率,具备良好泛化与鲁棒性。代码与数据集已开源。

原文摘要 · Abstract (English)

Recent advances in diffusion models have enabled the creation of deceptively real images, posing significant security risks when misused. In this study, we empirically show that different timesteps of DDIM inversion reveal varying subtle distinctions between synthetic and real images that are extractable for detection, in the forms of such as Fourier power spectrum high-frequency discrepancies and inter-pixel variance distributions. Based on these observations, we propose a novel synthetic image detection method that directly utilizes features of intermediately noised images by training an ensemble on multiple noised timesteps, circumventing conventional reconstruction-based strategies. To enhance human comprehension, we introduce a metric-grounded explanation generation and refinement module to identify and explain AI-generated flaws. Additionally, we construct the GenHard and GenExplain benchmarks to provide detection samples of greater difficulty and high-quality rationales for fake images. Extensive experiments show that our method achieves state-of-the-art performance with 98.91% and 95.89% detection accuracy on regular and challenging samples respectively, and demonstrates generalizability and robustness. Our code and datasets are available at https://github.com/Shadowlized/ESIDE.

图像检测可解释性扩散模型生成对抗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。