用视觉语音识别检测深度伪造,零样本泛化能力强
Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
- 基于预训练视觉语音识别特征,提取视频时序特征区分真假
- 在零样本设置下优于现有方法,对六种生成技术均有效
- 提出新数据集Authentica-Vox/HDTF,适合对抗性检测研究
深度伪造生成技术发展迅速,已能生成高度逼真的图像、视频和音频,引发滥用担忧。为此,亟需鲁棒可靠的深度伪造检测方法。本文提出FauxNet网络,基于预训练的视觉语音识别(VSR)特征,从视频中提取时序特征以识别真实与伪造内容。研究聚焦于零样本检测——即具备泛化能力的检测,FauxNet在此设置下持续超越当前最优方法。此外,FauxNet可溯源生成技术。本文构建了两个新数据集:Authentica-Vox和Authentica-HDTF,共包含约38,000个真实与伪造视频,后者涵盖六种近期深度伪造生成技术。在Authentica与FaceForensics++数据集上进行了广泛分析与实验,验证了FauxNet的优势。所建数据集将公开共享。
原文摘要 · Abstract (English)
Deepfake generation has witnessed remarkable progress, contributing to highly realistic generated images, videos, and audio. While technically intriguing, such progress has raised serious concerns related to the misuse of manipulated media. To mitigate such misuse, robust and reliable deepfake detection is urgently needed. Towards this, we propose a novel network FauxNet, which is based on pre-trained Visual Speech Recognition (VSR) features. By extracting temporal VSR features from videos, we identify and segregate real videos from manipulated ones. The holy grail in this context has to do with zero-shot detection, i.e., generalizable detection, which we focus on in this work. FauxNet consistently outperforms the state-of-the-art in this setting. In addition, FauxNet is able to attribute - distinguish between generation techniques from which the videos stem. Finally, we propose new datasets, referred to as Authentica-Vox and Authentica-HDTF, comprising about 38,000 real and fake videos in total, the latter created with six recent deepfake generation techniques. We provide extensive analysis and results on the Authentica datasets and FaceForensics++, demonstrating the superiority of FauxNet. The Authentica datasets will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。