arXiv:2603.05852cs.SD2026-03被引 1

测试现有语音伪造检测方法在真实场景下的泛化能力

How Well Do Current Speech Deepfake Detection Methods Generalize to the Real World?

  • 构建多语言真实场景数据集ML-ITW,覆盖14种语言和7大平台
  • 三种检测模型在真实环境下性能大幅下降,跨语言表现差
  • 适合关注语音安全、真实数据评估的研究者使用

语音合成与语音转换技术的进展显著提升了生成音频的自然度和真实性。与此同时,社交媒体平台上的编码、压缩和传输机制不断演进,进一步掩盖了深度伪造的痕迹。这些因素使真实环境中的可靠检测变得更加困难,凸显了建立代表性评估基准的必要性。为此,我们提出了ML-ITW(多语言野外)数据集,涵盖14种语言、7个主流平台以及180位公众人物,总音频时长达28.39小时。我们评估了三种检测范式:端到端神经模型、自监督特征方法(SSL)和音频大语言模型(Audio LLM)。实验结果表明,在不同语言和真实声学条件下,各类检测模型均出现显著性能下降,揭示了现有检测器在实际应用中泛化能力有限。ML-ITW数据集已公开可用。

原文摘要 · Abstract (English)

Recent advances in speech synthesis and voice conversion have greatly improved the naturalness and authenticity of generated audio. Meanwhile, evolving encoding, compression, and transmission mechanisms on social media platforms further obscure deepfake artifacts. These factors complicate reliable detection in real-world environments, underscoring the need for representative evaluation benchmarks. To this end, we introduce ML-ITW (Multilingual In-The-Wild), a multilingual dataset covering 14 languages, seven major platforms, and 180 public figures, totaling 28.39 hours of audio. We evaluate three detection paradigms: end-to-end neural models, self-supervised feature-based (SSL) methods, and audio large language models (Audio LLMs). Experimental results reveal significant performance degradation across diverse languages and real-world acoustic conditions, highlighting the limited generalization ability of existing detectors in practical scenarios. The ML-ITW dataset is publicly available.

语音伪造真实场景多语言检测评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。