首个多语言音视频深度伪造检测开放集基准,覆盖8种语言超250小时数据。
MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark
- 构建8语言音视频数据集,60%为生成伪造内容,涵盖7类生成模型。
- 在训练时仅暴露部分模型与语言,模拟真实世界开放集检测挑战。
- 现有顶尖检测器在此场景下性能显著下降,凸显实际应用难点。
我们提出了首个大规模多语言音视频深度伪造检测开放集基准。数据集包含超过250小时的真实与伪造视频,覆盖八种语言,其中60%的数据由生成模型创建。每种语言的伪造视频均使用七种不同深度伪造生成模型生成,选型基于生成内容质量。训练、验证与测试集的划分方式确保训练阶段仅提供部分生成模型和语言,从而创建多个具有挑战性的开放集评估场景。我们对近期文献中提出的多种预训练及微调深度伪造检测器进行了实验。结果表明,当前最先进的检测器在本开放集设置下无法维持其性能水平。数据与代码已公开发布于:https://huggingface.co/datasets/unibuc-cs/MAVOS-DD。
原文摘要 · Abstract (English)
We present the first large-scale open-set benchmark for multilingual audio-video deepfake detection. Our dataset comprises over 250 hours of real and fake videos across eight languages, with 60% of data being generated. For each language, the fake videos are generated with seven distinct deepfake generation models, selected based on the quality of the generated content. We organize the training, validation and test splits such that only a subset of the chosen generative models and languages are available during training, thus creating several challenging open-set evaluation setups. We perform experiments with various pre-trained and fine-tuned deepfake detectors proposed in recent literature. Our results show that state-of-the-art detectors are not currently able to maintain their performance levels when tested in our open-set scenarios. We publicly release our data and code at: https://huggingface.co/datasets/unibuc-cs/MAVOS-DD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。