arXiv:2506.00462cs.SDcs.AI2025-06Conference of the …被引 4

构建跨领域多语言语音伪造检测基准,揭示现有方法泛化能力严重不足。

XMAD-Bench: Cross-Domain Multilingual Audio Deepfake Benchmark

  • 构建668.8小时真实与伪造语音数据集,训练与测试使用不同生成模型和说话人。
  • 同一模型在跨域测试中准确率降至随机水平,远低于99%的域内表现。
  • 适合研究鲁棒性语音伪造检测的学者,尤其关注实际应用中的泛化问题。

近年来音频生成技术发展迅速,导致深度伪造音频增多,公众面临金融诈骗、身份盗用和虚假信息的风险。尽管许多音频伪造检测方法在域内测试中报告准确率接近99%,但这些方法通常仅在同源生成模型下评估。为此,我们提出XMAD-Bench,一个大规模跨域多语言音频深度伪造基准,包含668.8小时的真实与伪造语音。该数据集中训练与测试集的说话人、生成方法及真实音频来源均不相同,形成具有挑战性的跨域评估场景,可模拟真实世界检测环境。我们的实验表明,同一模型在域内准确率可达100%,但在跨域测试中性能显著下降,有时仅相当于随机猜测。该基准凸显了开发具备跨语言、跨说话人、跨生成方法和跨数据源泛化能力的鲁棒检测器的迫切需求。数据集已公开于https://github.com/ristea/xmad-bench/。

原文摘要 · Abstract (English)

Recent advances in audio generation led to an increasing number of deepfakes, making the general public more vulnerable to financial scams, identity theft, and misinformation. Audio deepfake detectors promise to alleviate this issue, with many recent studies reporting accuracy rates close to 99%. However, these methods are typically tested in an in-domain setup, where the deepfake samples from the training and test sets are produced by the same generative models. To this end, we introduce XMAD-Bench, a large-scale cross-domain multilingual audio deepfake benchmark comprising 668.8 hours of real and deepfake speech. In our novel dataset, the speakers, the generative methods, and the real audio sources are distinct across training and test splits. This leads to a challenging cross-domain evaluation setup, where audio deepfake detectors can be tested "in the wild". Our in-domain and cross-domain experiments indicate a clear disparity between the in-domain performance of deepfake detectors, which is usually as high as 100%, and the cross-domain performance of the same models, which is sometimes similar to random chance. Our benchmark highlights the need for the development of robust audio deepfake detectors, which maintain their generalization capacity across different languages, speakers, generative methods, and data sources. Our benchmark is publicly released at https://github.com/ristea/xmad-bench/.

语音伪造检测基准跨域泛化多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。