arXiv:2512.21695cs.CV2025-12中稿 · publication in 202…

融合频谱与语义特征,提升生成图像检测鲁棒性

FUSE: Unifying Spectral and Semantic Cues for Robust AI-Generated Image Detection

  • 用傅里叶变换提取频谱特征,结合CLIP视觉编码器获取语义特征
  • 在GenImage等5个数据集上平均准确率达91.36%,Chameleon上达94.96%均值AP
  • 对高保真图像仍保持稳定检测能力,适合多生成器场景应用

生成模型的快速演进加剧了对可信AI生成图像检测的需求。为此,我们提出FUSE,一种融合快速傅里叶变换提取的频谱特征与CLIP视觉编码器获得的语义特征的混合系统。特征被融合为联合表示,并采用两阶段渐进式训练。在GenImage、WildFake、DiTFake、GPT-ImgEval和Chameleon数据集上的评估表明,该方法在多种生成器间具有强泛化能力。FUSE(第一阶段)在Chameleon基准上达到最先进性能,于GenImage数据集上实现91.36%平均准确率,所有测试生成器上平均准确率为88.71%,平均精度均值达94.96%。第二阶段训练进一步提升了多数生成器的性能。不同于现有方法在Chameleon高保真图像上表现不佳,本方法在多种生成器下均保持鲁棒性。结果凸显了融合频谱与语义特征在通用生成图像检测中的优势。

原文摘要 · Abstract (English)

The fast evolution of generative models has heightened the demand for reliable detection of AI-generated images. To tackle this challenge, we introduce FUSE, a hybrid system that combines spectral features extracted through Fast Fourier Transform with semantic features obtained from the CLIP's Vision encoder. The features are fused into a joint representation and trained progressively in two stages. Evaluations on GenImage, WildFake, DiTFake, GPT-ImgEval and Chameleon datasets demonstrate strong generalization across multiple generators. Our FUSE (Stage 1) model demonstrates state-of-the-art results on the Chameleon benchmark. It also attains 91.36% mean accuracy on the GenImage dataset, 88.71% accuracy across all tested generators, and a mean Average Precision of 94.96%. Stage 2 training further improves performance for most generators. Unlike existing methods, which often perform poorly on high-fidelity images in Chameleon, our approach maintains robustness across diverse generators. These findings highlight the benefits of integrating spectral and semantic features for generalized detection of images generated by AI.

图像检测生成对抗多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。