arXiv:2607.06354cs.CVcs.CR2026-07

通过增强的RGB噪声学习,提升跨模型伪造图像检测能力。

Generalized Synthetic Image Detection with Enhanced RGB-Noise Representation Learning

论文配图:Generalized Synthetic Image Detection with Enhanced RGB-Noise Representation Learning
图 1 · 摘自论文原文
  • 双分支结构融合语义与高频噪声特征,动态调制增强表示
  • 在8个公开数据集上达到最优性能,泛化性与鲁棒性显著提升
  • 适合需要高可靠伪造检测的应用场景,如内容审核

大规模生成模型的快速发展加速了高度欺骗性AI生成图像的传播,使得泛化性强的合成图像检测成为迫切需求。现有取证网络因依赖单一领域表示和传统二分类优化,在跨模型泛化和真实世界退化场景下表现不佳。为此,我们提出RNSIDNet,一种通过增强RGB-噪声表示学习实现鲁棒检测的新框架。该方法采用双分支架构:由注意力精炼的CLIP主干提取全局RGB语义,通过特征逐元素线性调制(FiLM)模块动态调控贝亚尔卷积捕捉的高频噪声伪影。为进一步优化表示,设计硬样本感知对比学习(HSCL)策略,显式惩罚困难样本,重塑潜在特征空间以最大化真实与合成域间的判别边界。在八个公开基准数据集上的广泛实验表明,本模型实现了领先性能,具备卓越的泛化能力、鲁棒性及计算效率。代码与数据集将公开于https://github.com/multimediaFor/RNSIDNet。

原文摘要 · Abstract (English)

The rapid advancement of large-scale generative models has accelerated the spread of highly deceptive AI-generated images, making generalized synthetic image detection a critical imperative. Existing forensic networks often struggle with cross-model generalization and realworld degradations due to their reliance on single-domain representations and conventional binary classification optimization. To overcome these limitations, we propose RNSIDNet, a novel forensic framework that achieves robust detection through enhanced RGB-Noise representation learning. Specifically, our method employs a dual-branch architecture where global RGB semantics, extracted by an attention-refined CLIP backbone, dynamically modulate highfrequency noise artifacts captured by Bayar convolutions via a Feature-wise Linear Modulation (FiLM) module. To further enhance the learned representations, we design a Hard Sample-aware Contrastive Learning (HSCL) strategy. By explicitly penalizing challenging training samples, HSCL reshapes the latent feature space to maximize the discriminative margin between pristine and synthetic domains. Extensive experiments across eight public benchmark datasets verify that our model achieves state-of-the-art performance, delivering superior generalization ability, robustness, and computational efficiency. Code and dataset will be publicly available on https://github.com/multimediaFor/RNSIDNet.

图像取证伪造检测噪声分析对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。