arXiv:2603.10598cs.CV2026-03中稿 · CVPR被引 3

通过分析图像隐空间层间一致性,提升合成图像检测的泛化能力。

Layer Consistency Matters: Elegant Latent Transition Discrepancy for Generalizable Synthetic Image Detection

  • 基于隐空间层间特征转移的一致性差异设计检测方法
  • 在三个数据集上平均准确率比基线高14.35%
  • 对未见过的生成模型和类型具有强鲁棒性,适合安全验证场景

生成模型的快速发展显著提升了AI合成图像的真实感与可及性。然而,其高度逼真的特性使合成图像越来越难以与真实照片区分,带来媒体可信度和内容篡改等严重安全风险。尽管已有大量研究致力于合成图像检测,但多数方法依赖特定模型的痕迹或低层统计特征,泛化能力差。本文发现:真实图像在隐空间中保持语义注意力与结构一致性,层间特征转移更稳定;而合成图像则表现出明显差异。为此,提出一种新方法——隐空间转移差异(LTD),通过自适应识别最具判别力的网络层,量化层间转移差异。得益于层间判别建模,该方法在包含多种GAN和扩散模型的三个数据集上,平均准确率比基线提升14.35%。大量实验表明,LTD优于当前先进方法,在检测精度、泛化性和鲁棒性上均表现优异。代码已开源。

原文摘要 · Abstract (English)

Recent rapid advancement of generative models has significantly improved the fidelity and accessibility of AI-generated synthetic images. While enabling various innovative applications, the unprecedented realism of these synthetics makes them increasingly indistinguishable from authentic photographs, posing serious security risks, such as media credibility and content manipulation. Although extensive efforts have been dedicated to detecting synthetic images, most existing approaches suffer from poor generalization to unseen data due to their reliance on model-specific artifacts or low-level statistical cues. In this work, we identify a previously unexplored distinction that real images maintain consistent semantic attention and structural coherence in their latent representations, exhibiting more stable feature transitions across network layers, whereas synthetic ones present discernible distinct patterns. Therefore, we propose a novel approach termed latent transition discrepancy (LTD), which captures the inter-layer consistency differences of real and synthetic images. LTD adaptively identifies the most discriminative layers and assesses the transition discrepancies across layers. Benefiting from the proposed inter-layer discriminative modeling, our approach exceeds the base model by 14.35\% in mean Acc across three datasets containing diverse GANs and DMs. Extensive experiments demonstrate that LTD outperforms recent state-of-the-art methods, achieving superior detection accuracy, generalizability, and robustness. The code is available at https://github.com/yywencs/LTD

图像检测生成模型泛化能力隐空间分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。