通过边缘差异分析提升ViT对生成图像的检测精度
Edge-Enhanced Vision Transformer Framework for Accurate AI-Generated Image Detection
- 融合ViT与新型边缘处理模块,捕捉生成图像的平滑纹理特征
- 在CIFAKE数据集上达97.75%准确率和97.77%F1分数
- 轻量高效、可解释性强,适合真实场景内容验证
生成模型的快速发展导致高度逼真的AI生成图像日益泛滥,给数字取证与内容认证带来严峻挑战。传统检测方法依赖提取全局特征的深度学习模型,常忽略细微结构不一致,且计算开销大。为此,我们提出一种混合检测框架,结合微调的视觉变换器(ViT)与新颖的基于边缘的图像处理模块。该模块通过计算平滑前后边缘差图的方差,利用生成图像通常具有更平滑纹理、较弱边缘和更低噪声的特性。作为ViT预测的后处理步骤,该模块提升了对细粒度结构线索的敏感性,同时保持高效计算。在CIFAKE、Artistic和Custom Curated数据集上的大量实验表明,所提框架在所有基准上表现优异,在CIFAKE上达到97.75%准确率和97.77%F1分数,超越广泛采用的先进模型。结果证明该方法是轻量、可解释且高效的解决方案,适用于静态图像与视频帧,极具实际应用价值。
原文摘要 · Abstract (English)
The rapid advancement of generative models has led to a growing prevalence of highly realistic AI-generated images, posing significant challenges for digital forensics and content authentication. Conventional detection methods mainly rely on deep learning models that extract global features, which often overlook subtle structural inconsistencies and demand substantial computational resources. To address these limitations, we propose a hybrid detection framework that combines a fine-tuned Vision Transformer (ViT) with a novel edge-based image processing module. The edge-based module computes variance from edge-difference maps generated before and after smoothing, exploiting the observation that AI-generated images typically exhibit smoother textures, weaker edges, and reduced noise compared to real images. When applied as a post-processing step on ViT predictions, this module enhances sensitivity to fine-grained structural cues while maintaining computational efficiency. Extensive experiments on the CIFAKE, Artistic, and Custom Curated datasets demonstrate that the proposed framework achieves superior detection performance across all benchmarks, attaining 97.75% accuracy and a 97.77% F1-score on CIFAKE, surpassing widely adopted state-of-the-art models. These results establish the proposed method as a lightweight, interpretable, and effective solution for both still images and video frames, making it highly suitable for real-world applications in automated content verification and digital forensics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。