arXiv:2507.03054cs.CVcs.AI2025-07被引 4

通过分析扩散过程中的潜在空间轨迹,提升生成图像检测精度。

LATTE: Latent Trajectory Embedding for Diffusion-Generated Image Detection

  • 建模多步去噪中潜在表示的演化轨迹,捕捉生成痕迹
  • 在跨生成器和跨数据集场景下性能超越现有方法
  • 适合需要高鲁棒性检测能力的研究与应用

基于扩散模型的图像生成技术快速发展,使真实与生成图像的区分愈发困难,威胁数字媒体可信度。现有检测方法多依赖单步重建误差,忽视去噪过程的时序特性。本文提出LATTE——潜空间轨迹嵌入方法,通过建模多个去噪步骤中潜在表示的演化路径,揭示真实与生成图像间的细微差异模式。在GenImage、Chameleon和Diffusion Forensics等多个基准测试中,LATTE表现优异,尤其在跨生成器和跨数据集场景下优势显著,验证了潜空间轨迹建模的有效性。代码已开源。

原文摘要 · Abstract (English)

The rapid advancement of diffusion-based image generators has made it increasingly difficult to distinguish generated from real images. This erodes trust in digital media, making it critical to develop generated image detectors that remain reliable across different generators. While recent approaches leverage diffusion denoising cues, they typically rely on single-step reconstruction errors and overlook the sequential nature of the denoising process. In this work, we propose LATTE - LATent Trajectory Embedding - a novel approach that models the evolution of latent embeddings across multiple denoising steps. Instead of treating each denoising step in isolation, LATTE captures the trajectory of these representations, revealing subtle and discriminative patterns that distinguish real from generated images. Experiments on several benchmarks, such as GenImage, Chameleon, and Diffusion Forensics, show that LATTE achieves superior performance, especially in challenging cross-generator and cross-dataset scenarios, highlighting the potential of latent trajectory modeling. The code is available on the following link: https://github.com/AnaMVasilcoiu/LATTE-Diffusion-Detector.

图像检测扩散模型潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。