融合CNN与ViT提升AI生成图像识别准确率
AI-Generated Image Recognition via Fusion of CNNs and Vision Transformers

- 结合CNN与视觉Transformer的多模型融合策略
- 在CIFAKE数据集上达到97.32%识别准确率
- 适合关注AI图像检测与数据可信验证的研究者
近年来,合成数据技术的进步使得高质量图像生成成为可能,模糊了真实图像与AI生成图像之间的界限。这一趋势对数据可靠性与真实性构成严峻挑战,亟需可靠的检测方法。本文提出一种基于多检测方法融合的鲁棒识别方案,充分利用CNN与视觉Transformer的优势。在CIFAKE数据集上的大量实验表明,该模型表现优异,识别准确率达到97.32%。结果证明该方法能有效区分AI生成图像与真实图像,为合成数据泛滥背景下的数据认证技术发展提供了有力支持。
原文摘要 · Abstract (English)
Recent advancements in synthetic data technology have opened a new era where images of remarkable quality are generated, blurring the lines between real-life images and those produced by Artificial Intelligence (AI). This evolution poses a significant challenge to ensuring the reliability and authenticity of data, underscoring the need for robust detection methods. In this paper, we present a robust approach aimed at addressing these pressing concerns. Our methodology revolves around leveraging fusion strategies, combining the strengths of multiple detection methods for identifying AI-generated images. Through extensive experimentation on the CIFAKE dataset, our model showcases remarkable performance, achieving an impressive accuracy rate of 97.32%. This accomplishment underscores the efficacy of our approach in accurately distinguishing between AI-generated images and real-life images, thus contributing to the advancement of data authentication techniques amidst the proliferation of synthetic data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。