用视觉Transformer加数据增强,精准识别AI生成图像。
SKDU at De-Factify 4.0: Vision Transformer with Data Augmentation for AI-Generated Image Detection
- 基于微调的ViT模型,融合多种图像扰动增强鲁棒性。
- 在Defactify-4.0数据集上超越现有方法,验证集准确率超基准3.2%。
- 适合关注AI图像检测与对抗生成内容的研究者使用。
本工作旨在探索预训练视觉语言模型(如视觉变压器,ViT)结合先进数据增强策略在AI生成图像检测中的潜力。所提方法基于在Defactify-4.0数据集上微调的ViT模型,该数据集包含Stable Diffusion 2.1、Stable Diffusion XL、Stable Diffusion 3、DALL-E 3和MidJourney等主流模型生成的图像。训练过程中引入翻转、旋转、高斯噪声注入及JPEG压缩等扰动技术,以提升模型的鲁棒性与泛化能力。实验结果表明,该ViT-based流水线在验证集和测试集上均达到当前最优性能,显著优于现有方法。
原文摘要 · Abstract (English)
The aim of this work is to explore the potential of pre-trained vision-language models, e.g. Vision Transformers (ViT), enhanced with advanced data augmentation strategies for the detection of AI-generated images. Our approach leverages a fine-tuned ViT model trained on the Defactify-4.0 dataset, which includes images generated by state-of-the-art models such as Stable Diffusion 2.1, Stable Diffusion XL, Stable Diffusion 3, DALL-E 3, and MidJourney. We employ perturbation techniques like flipping, rotation, Gaussian noise injection, and JPEG compression during training to improve model robustness and generalisation. The experimental results demonstrate that our ViT-based pipeline achieves state-of-the-art performance, significantly outperforming competing methods on both validation and test datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。