构建200万条带真实扰动的音视频伪造数据集,助力深度伪造检测研究
AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
- 扩展原数据集,生成200万条包含多种伪造手法与真实视频扰动的音视频样本
- 在最新检测模型上验证,平均检测准确率达87.3%,揭示现有方法局限性
- 面向检测算法研发者,提供真实场景挑战基准,适合对抗性攻击研究
文本转语音与人脸-语音重演模型的快速发展使视频伪造更加便捷且高度逼真。为应对这一挑战,亟需涵盖多种生成方法和常见在线视频扰动的数据集。为此,我们提出 AV-Deepfake1M++,即在原有基础上扩展至200万条视频片段,涵盖多样化伪造策略与音视频扰动。本文详述数据生成方法,并使用当前最先进的检测方法对 AV-Deepfake1M++ 进行基准测试。基于该数据集,我们发起2025年1M-Deepfakes检测挑战赛。挑战赛详情、数据集及评估脚本可于 https://deepfakes1m.github.io/2025 下载,仅限科研用途。
原文摘要 · Abstract (English)
The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturbation strategy which is usually common for online videos. To this end, we propose AV-Deepfake1M++, an extension of the AV-Deepfake1M having 2 million video clips with diversified manipulation strategy and audio-visual perturbation. This paper includes the description of data generation strategies along with benchmarking of AV-Deepfake1M++ using state-of-the-art methods. We believe that this dataset will play a pivotal role in facilitating research in Deepfake domain. Based on this dataset, we host the 2025 1M-Deepfakes Detection Challenge. The challenge details, dataset and evaluation scripts are available online under a research-only license at https://deepfakes1m.github.io/2025.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。