通过像素级扰动提升AI生成图像检测的泛化能力
Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image Detection
- 在检测前对图像进行像素值分布扰乱,打破依赖特定语义线索的缺陷
- 在多种生成模型上显著提升检测器跨模型性能,最高提升18.7%准确率
- 适合需要跨模型通用检测能力的研究者和实际应用部署
生成技术的快速发展亟需可靠的AI生成图像检测方法。当前检测器普遍存在泛化能力差的问题,因其常过度依赖特定生成模型的语义特征,而非学习普遍存在的生成痕迹。为此,我们提出一种简单而有效的像素级映射预处理步骤,通过扰乱图像的像素值分布,破坏检测器常用的非必要语义模式。这迫使检测器关注图像生成过程中更基础、更具普适性的高频特征。在基于GAN与扩散模型的多种生成器上进行的全面实验表明,该方法显著提升了主流检测器的跨生成器性能。进一步分析验证了:破坏语义线索是实现泛化的核心机制。
原文摘要 · Abstract (English)
The rapid evolution of generative technologies necessitates reliable methods for detecting AI-generated images. A critical limitation of current detectors is their failure to generalize to images from unseen generative models, as they often overfit to source-specific semantic cues rather than learning universal generative artifacts. To overcome this, we introduce a simple yet remarkably effective pixel-level mapping pre-processing step to disrupt the pixel value distribution of images and break the fragile, non-essential semantic patterns that detectors commonly exploit as shortcuts. This forces the detector to focus on more fundamental and generalizable high-frequency traces inherent to the image generation process. Through comprehensive experiments on GAN and diffusion-based generators, we show that our approach significantly boosts the cross-generator performance of state-of-the-art detectors. Extensive analysis further verifies our hypothesis that the disruption of semantic cues is the key to generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。