用多层视觉特征提升AI生成图像检测效果
Rethinking the Use of Vision Transformers for AI-Generated Image Detection
- 从ViT多层特征中动态融合信息,改进检测性能
- 早期层特征比末层更准,跨模型泛化能力更强
- 方法可推广至DINOv2等其他预训练模型
CLIP-ViT提取的丰富特征已被广泛用于AI生成图像检测。现有方法主要依赖最后一层特征,我们系统分析了各层特征的贡献。研究发现,早期层提供更局部且更具泛化性的特征,在检测任务中常优于末层特征。不同层捕捉数据的不同方面,各自对检测有独特贡献。受此启发,我们提出一种新型自适应方法MoLD,通过门控机制动态融合多层特征。在GAN与扩散模型生成图像上的大量实验表明,MoLD显著提升检测性能,增强对多种生成模型的泛化能力,并在真实场景中表现出强鲁棒性。最后,我们展示了该方法在DINOv2等其他预训练ViT上的可扩展性和通用性。
原文摘要 · Abstract (English)
Rich feature representations derived from CLIP-ViT have been widely utilized in AI-generated image detection. While most existing methods primarily leverage features from the final layer, we systematically analyze the contributions of layer-wise features to this task. Our study reveals that earlier layers provide more localized and generalizable features, often surpassing the performance of final-layer features in detection tasks. Moreover, we find that different layers capture distinct aspects of the data, each contributing uniquely to AI-generated image detection. Motivated by these findings, we introduce a novel adaptive method, termed MoLD, which dynamically integrates features from multiple ViT layers using a gating-based mechanism. Extensive experiments on both GAN- and diffusion-generated images demonstrate that MoLD significantly improves detection performance, enhances generalization across diverse generative models, and exhibits robustness in real-world scenarios. Finally, we illustrate the scalability and versatility of our approach by successfully applying it to other pre-trained ViTs, such as DINOv2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。