arXiv:2508.09487cs.CV2025-08被引 2

通过语义差异检测生成图像,跨模型泛化能力强。

Semantic-Aware Reconstruction Error for Detecting AI-Generated Images

  • 用图文重建的语义差异衡量真实与生成图像区别
  • 在GenImage和ForenSynths上优于现有方法
  • 适合需要跨模型检测的AI图像安全场景

近年来,随着图像生成技术快速发展,AI生成图像的检测受到广泛关注。现有方法多依赖特定模型的伪影,面对未见的、分布外(OOD)生成模型时性能显著下降。为此,本文提出一种新表征——语义感知重建误差(SARE),通过度量图像与其标题引导重建之间的语义差异来检测假图像。核心假设是:真实图像的标题常无法完整描述复杂视觉内容,重建过程中会产生明显语义偏移;而生成图像与标题高度一致,语义变化极小。通过量化这种语义偏移,SARE能提供对多种生成模型都鲁棒且具有判别力的特征。此外,设计了一种融合模块,利用交叉注意力机制将SARE嵌入主干检测器,使图像特征自适应关注来自SARE的语义表示。实验表明,该方法在GenImage和ForenSynths等基准上均超越现有基线。通过详细分析语义偏移,进一步验证了标题引导的有效性,可显著提升检测鲁棒性。

原文摘要 · Abstract (English)

Recently, AI-generated image detection has gained increasing attention, as the rapid advancement of image generation technologies has raised serious concerns about their potential misuse. While existing detection methods have achieved promising results, their performance often degrades significantly when facing fake images from unseen, out-of-distribution (OOD) generative models, since they primarily rely on model-specific artifacts and thus overfit to the models used for training. To address this limitation, we propose a novel representation, namely Semantic-Aware Reconstruction Error (SARE), that measures the semantic difference between an image and its caption-guided reconstruction. The key hypothesis behind SARE is that real images, whose captions often fail to fully capture their complex visual content, may undergo noticeable semantic shifts during the caption-guided reconstruction process. In contrast, fake images, which closely align with their captions, show minimal semantic changes. By quantifying these semantic shifts, SARE provides a robust and discriminative feature for detecting fake images across diverse generative models. Additionally, we introduce a fusion module that integrates SARE into the backbone detector via a cross-attention mechanism. Image features attend to semantic representations extracted from SARE, enabling the model to adaptively leverage semantic information. Experimental results demonstrate that the proposed method achieves strong generalization, outperforming existing baselines on benchmarks including GenImage and ForenSynths. We further validate the effectiveness of caption guidance through a detailed analysis of semantic shifts, confirming its ability to enhance detection robustness.

图像检测生成模型语义分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。