构建首个多模态深度伪造检测统一基准,助力对抗虚假音视频威胁。
DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection
- 整合21种伪造流程,构建超大规模多模态数据集Mega-MMDF
- 包含110万伪造样本与10万真实样本,覆盖多种伪造技术组合
- 提供标准化评测平台,适合研究者验证新方法或对比现有模型
生成式AI的滥用导致大量伪造音视频内容泛滥,严重威胁社会安全(如金融诈骗、社会动荡)。尽管已有初步应对研究,但缺乏充足多样训练数据及标准化评估基准,限制了深入探索。为此,我们首先构建了Mega-MMDF——一个大规模、多样化且高质量的多模态深度伪造检测数据集。该数据集结合10种音频伪造方法、12种视觉伪造方法和6种音频驱动人脸重演技术,共生成21种伪造流程,目前包含10万真实样本与110万伪造样本,是当前最全面的多模态伪造数据集之一,并支持持续扩展。基于此,我们提出DeepfakeBench-MM,首个统一的多模态深度伪造检测基准。它建立贯穿检测全流程的标准协议,可作为评估现有方法与探索新策略的通用平台,目前支持5个数据集和11种多模态检测器。通过全面评估与深入分析,我们在数据增强、复合伪造等角度发现多个关键规律。我们认为,DeepfakeBench-MM与大规模的Mega-MMDF将共同成为推进多模态深度伪造检测的基础设施。
原文摘要 · Abstract (English)
The misuse of advanced generative AI models has resulted in the widespread proliferation of falsified data, particularly forged human-centric audiovisual content, which poses substantial societal risks (e.g., financial fraud and social instability). In response to this growing threat, several works have preliminarily explored countermeasures. However, the lack of sufficient and diverse training data, along with the absence of a standardized benchmark, hinder deeper exploration. To address this challenge, we first build Mega-MMDF, a large-scale, diverse, and high-quality dataset for multimodal deepfake detection. Specifically, we employ 21 forgery pipelines through the combination of 10 audio forgery methods, 12 visual forgery methods, and 6 audio-driven face reenactment methods. Mega-MMDF currently contains 0.1 million real samples and 1.1 million forged samples, making it one of the largest and most diverse multimodal deepfake datasets, with plans for continuous expansion. Building on it, we present DeepfakeBench-MM, the first unified benchmark for multimodal deepfake detection. It establishes standardized protocols across the entire detection pipeline and serves as a versatile platform for evaluating existing methods as well as exploring novel approaches. DeepfakeBench-MM currently supports 5 datasets and 11 multimodal deepfake detectors. Furthermore, our comprehensive evaluations and in-depth analyses uncover several key findings from multiple perspectives (e.g., augmentation, stacked forgery). We believe that DeepfakeBench-MM, together with our large-scale Mega-MMDF, will serve as foundational infrastructures for advancing multimodal deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。