融合多种视觉线索,提升生成图像检测的跨模型泛化能力。
Aggregating Diverse Cue Experts for AI-Generated Image Detection
- 设计统一框架,动态融合空间、频域和色彩不一致线索。
- 在八种生成器上平均准确率领先7.4%,突破现有方法极限。
- 适合需要强泛化能力的图像真实性验证场景。
图像生成模型的快速发展给AI生成图像检测带来了挑战。现有方法多依赖特定模型特征,导致过拟合和泛化能力差。本文提出多线索聚合网络(MCAN),通过混合编码器适配器动态处理多种互补线索,实现更自适应的特征表示。所用线索包括输入图像内容、高频边缘细节,以及新引入的色度不一致(CI)线索——通过对强度值归一化,突出真实图像采集过程中的噪声特征,使其与生成内容的噪声模式更易区分。不同于以往方法,MCAN在统一框架中融合空间、频域与色度信息,使线索更具真实性判别力,显著提升跨模型泛化性能。在GenImage、Chameleon和UniversalFakeDetect三个基准测试上验证,其在GenImage数据集上对八种不同生成器的平均准确率比最优现有方法最高提升7.4%。
原文摘要 · Abstract (English)
The rapid emergence of image synthesis models poses challenges to the generalization of AI-generated image detectors. However, existing methods often rely on model-specific features, leading to overfitting and poor generalization. In this paper, we introduce the Multi-Cue Aggregation Network (MCAN), a novel framework that integrates different yet complementary cues in a unified network. MCAN employs a mixture-of-encoders adapter to dynamically process these cues, enabling more adaptive and robust feature representation. Our cues include the input image itself, which represents the overall content, and high-frequency components that emphasize edge details. Additionally, we introduce a Chromatic Inconsistency (CI) cue, which normalizes intensity values and captures noise information introduced during the image acquisition process in real images, making these noise patterns more distinguishable from those in AI-generated content. Unlike prior methods, MCAN's novelty lies in its unified multi-cue aggregation framework, which integrates spatial, frequency-domain, and chromaticity-based information for enhanced representation learning. These cues are intrinsically more indicative of real images, enhancing cross-model generalization. Extensive experiments on the GenImage, Chameleon, and UniversalFakeDetect benchmark validate the state-of-the-art performance of MCAN. In the GenImage dataset, MCAN outperforms the best state-of-the-art method by up to 7.4% in average ACC across eight different image generators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。