arXiv:2504.00463cs.CV2025-04被引 2

融合多种低层信息,提升AI生成图像检测的泛化能力

Exploring the Collaborative Advantage of Low-level Information on Generalizable AI-Generated Image Detection

  • 设计自适应专家注入框架,动态融合多种低层特征
  • 仅用4类ProGAN数据微调,即可在多类未见生成模型上达到领先性能
  • 适合需要高泛化能力的图像真实性检测场景

现有最先进的AI生成图像检测方法大多依赖从RGB图像中提取低层信息(如噪声模式)以提升泛化能力,但通常只关注单一类型,导致效果受限。通过实证分析发现,不同低层信息对不同类型的伪造具有不同的泛化能力,且简单的特征融合策略无法充分挖掘各层次信息的优势。为此,我们提出自适应低层专家注入(ALEI)框架:引入LoRA专家,使基于高层语义RGB图像训练的主干网络能够学习并整合多种低层信息;采用交叉注意力机制在中间层自适应融合特征;设计低层信息适配器,防止主干网络在后期建模中丢失低层特征建模能力;提出动态特征选择机制,根据当前图像动态选取最优检测特征,最大化泛化性能。大量实验表明,本方法仅在四类主流ProGAN数据上微调,即在包含未见GAN与扩散模型的多个数据集上表现优异,达到当前最佳水平。

原文摘要 · Abstract (English)

Existing state-of-the-art AI-Generated image detection methods mostly consider extracting low-level information from RGB images to help improve the generalization of AI-Generated image detection, such as noise patterns. However, these methods often consider only a single type of low-level information, which may lead to suboptimal generalization. Through empirical analysis, we have discovered a key insight: different low-level information often exhibits generalization capabilities for different types of forgeries. Furthermore, we found that simple fusion strategies are insufficient to leverage the detection advantages of each low-level and high-level information for various forgery types. Therefore, we propose the Adaptive Low-level Experts Injection (ALEI) framework. Our approach introduces Lora Experts, enabling the backbone network, which is trained with high-level semantic RGB images, to accept and learn knowledge from different low-level information. We utilize a cross-attention method to adaptively fuse these features at intermediate layers. To prevent the backbone network from losing the modeling capabilities of different low-level features during the later stages of modeling, we developed a Low-level Information Adapter that interacts with the features extracted by the backbone network. Finally, we propose Dynamic Feature Selection, which dynamically selects the most suitable features for detecting the current image to maximize generalization detection capability. Extensive experiments demonstrate that our method, finetuned on only four categories of mainstream ProGAN data, performs excellently and achieves state-of-the-art results on multiple datasets containing unseen GAN and Diffusion methods.

图像检测生成对抗泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。