arXiv:2412.17924cs.SDeess.AS2024-12被引 7

测试了英文训练的语音伪造检测模型在多语言场景下的表现

Are audio DeepFake detection models polyglots?

  • 用英文数据集训练的模型跨语言检测能力差
  • 目标语言数据缺失会显著降低检测效果
  • 适合研究多语言语音安全的学者参考

由于大多数语音深度伪造(DF)检测方法基于以英语为主的数据集训练,其在非英语语言中的适用性仍缺乏研究。本文构建了多语言语音伪造检测基准,评估了多种适配策略。实验聚焦于英语基准数据集训练的模型,以及同语言和跨语言适配方法。结果表明检测效能存在显著差异,凸显多语言场景的挑战。我们发现仅使用英语数据会削弱检测性能,强调目标语言数据的重要性。

原文摘要 · Abstract (English)

Since the majority of audio DeepFake (DF) detection methods are trained on English-centric datasets, their applicability to non-English languages remains largely unexplored. In this work, we present a benchmark for the multilingual audio DF detection challenge by evaluating various adaptation strategies. Our experiments focus on analyzing models trained on English benchmark datasets, as well as intra-linguistic (same-language) and cross-linguistic adaptation approaches. Our results indicate considerable variations in detection efficacy, highlighting the difficulties of multilingual settings. We show that limiting the dataset to English negatively impacts the efficacy, while stressing the importance of the data in the target language.

语音伪造多语言检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。