用视觉变压器集成检测深度伪造,通用性强且性能领先。
Towards Generalizable Deepfake Image Detection with Vision Transformers

- 采用DINOv2、AIMv2等多模型微调集成方法
- 在DF-Wild数据集上达AUC 96.77%、EER 9%
- 适合需要高泛化能力的深度伪造检测场景
当前深度伪造图像检测面临生成模型快速演进与现有方法泛化能力差的双重挑战。本文采用DINOv2、AIMv2及OpenCLIP的ViT-L/14等视觉变压器的集成方法,构建具备强泛化能力的检测模型。实验基于IEEE SP Cup 2025发布的DF-Wild数据集,该数据集涵盖多样且复杂的篡改手法与生成技术。初始实验使用基于空间特征的CNN分类器。结果表明,该集成方法优于单个模型和强基线CNN,在DF-Wild测试集上达到AUC 96.77%、Equal Error Rate(EER)仅9%,相较当前最优算法Effort在AUC和EER上分别提升7.05%和8%。该方案为SP Cup冠军解法,已提交ICASSP 2025。
原文摘要 · Abstract (English)
In today's day and age, we face a challenge in detecting deepfake images because of the fast evolution of modern generative models and the poor generalization capability of existing methods. In this paper, we use an ensemble of fine-tuned vision transformers like DINOv2, AIMv2 and OpenCLIP's ViT-L/14 to create generalizable method to detect deepfakes. We use the DF-Wild dataset released as part of the IEEE SP Cup 2025, because it uses a challenging and diverse set of manipulations and generation techniques. We started our experiments with CNN classifiers trained on spatial features. Experimental results show that our ensemble outperforms individual models and strong CNN baselines, achieving an AUC of 96.77% and an Equal Error Rate (EER) of just 9% on the DF-Wild test set, beating the state-of-the-art deepfake detection algorithm Effort by 7.05% and 8% in AUC and EER respectively. This was the winning solution for SP Cup, presented at ICASSP 2025.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。