通过可控噪声提升AI生成图像检测的泛化能力
How Noise Benefits AI-generated Image Detection
- 在特征空间注入正向激励噪声,抑制模型对捷径特征的依赖
- 在42种生成模型上平均准确率提升5.4%,达到新最优
- 适合需要强泛化能力的AI图像溯源与内容审核场景
生成模型的快速发展使得真实图像与合成图像越来越难以区分。尽管已有大量研究致力于检测AI生成图像,但跨分布泛化能力仍是长期挑战。我们发现该问题源于训练过程中对虚假捷径的依赖,并观察到微小的特征空间扰动可缓解捷径主导现象。为此,我们提出正向激励噪声的CLIP框架(PiN-CLIP),联合训练噪声生成器与检测网络,基于变分正向激励原则。具体地,通过视觉与语义特征的交叉注意力融合构建特征空间中的正向激励噪声。优化过程中,噪声被注入特征空间以微调视觉编码器,抑制对捷径敏感的方向,同时增强稳定的伪造线索,从而提取更具鲁棒性和泛化性的伪影表征。在包含42种不同生成模型的开放世界数据集上进行对比实验,本方法在平均准确率上相比现有方法提升5.4个百分点,达到新最佳性能。
原文摘要 · Abstract (English)
The rapid advancement of generative models has made real and synthetic images increasingly indistinguishable. Although extensive efforts have been devoted to detecting AI-generated images, out-of-distribution generalization remains a persistent challenge. We trace this weakness to spurious shortcuts exploited during training and we also observe that small feature-space perturbations can mitigate shortcut dominance. To address this problem in a more controllable manner, we propose the Positive-Incentive Noise for CLIP (PiN-CLIP), which jointly trains a noise generator and a detection network under a variational positive-incentive principle. Specifically, we construct positive-incentive noise in the feature space via cross-attention fusion of visual and categorical semantic features. During optimization, the noise is injected into the feature space to fine-tune the visual encoder, suppressing shortcut-sensitive directions while amplifying stable forensic cues, thereby enabling the extraction of more robust and generalized artifact representations. Comparative experiments are conducted on an open-world dataset comprising synthetic images generated by 42 distinct generative models. Our method achieves new state-of-the-art performance, with notable improvements of 5.4 in average accuracy over existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。