用视觉语言模型零样本检测深度伪造,效果优于传统方法。
Visual Language Models as Zero-Shot Deepfake Detectors
- 利用视觉语言模型的零样本能力进行深度伪造检测。
- 在6万张图像的数据集上表现超越多数现有方法。
- 适合关注高效、无需微调的深度伪造检测研究者。
当前利用GAN或扩散模型进行人脸替换的深度伪造现象,在数字媒体、身份验证等系统中构成严重且持续演化的威胁。现有大多数检测方法依赖于训练专用分类器来区分真实与篡改图像,仅聚焦图像域,未引入可增强鲁棒性的辅助任务。本文受视觉语言模型零样本能力启发,提出一种基于VLM的图像分类新方法,并评估其在深度伪造检测中的表现。具体而言,我们使用一个包含60,000张图像的高质量深度伪造数据集,结果显示我们的零样本模型性能优于几乎所有现有方法。随后,我们在主流数据集DFDC-P上对比了表现最佳的InstructBLIP架构在零样本和领域内微调两种场景下的表现,结果表明视觉语言模型显著优于传统分类器。
原文摘要 · Abstract (English)
The contemporary phenomenon of deepfakes, utilizing GAN or diffusion models for face swapping, presents a substantial and evolving threat in digital media, identity verification, and a multitude of other systems. The majority of existing methods for detecting deepfakes rely on training specialized classifiers to distinguish between genuine and manipulated images, focusing only on the image domain without incorporating any auxiliary tasks that could enhance robustness. In this paper, inspired by the zero-shot capabilities of Vision Language Models, we propose a novel VLM-based approach to image classification and then evaluate it for deepfake detection. Specifically, we utilize a new high-quality deepfake dataset comprising 60,000 images, on which our zero-shot models demonstrate superior performance to almost all existing methods. Subsequently, we compare the performance of the best-performing architecture, InstructBLIP, on the popular deepfake dataset DFDC-P against traditional methods in two scenarios: zero-shot and in-domain fine-tuning. Our results demonstrate the superiority of VLMs over traditional classifiers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。