arXiv:2409.06348cs.SDcs.AI2024-09被引 13

构建首个中英双语深伪语音检测基准,揭示现有方法在真实场景下表现大幅下降。

VoiceWukong: Benchmarking Deepfake Voice Detection

  • 收集34种工具生成的26.5万条英文与14.8万条中文伪造语音样本
  • 12个主流检测器平均误报率超20%,最佳模型AASIST2仅达13.5%的等错误率
  • 真人识别能力随欺骗程度变化,大模型对伪造语音无识别能力

随着文本转语音(TTS)和语音转换(VC)技术的快速发展,深伪语音检测变得日益重要。然而,学术界与工业界缺乏全面且直观的评测基准。现有数据集在语言多样性上有限,且未涵盖真实生产环境中的多种篡改手法。为此,我们提出VoiceWukong,一个用于评估深伪语音检测器性能的基准。我们首先收集了19款知名商业工具和15款开源工具生成的深伪语音,共创建38种数据变体,涵盖六类篡改类型,构建出包含265,200条英文与148,200条中文深伪语音样本的评测数据集。基于此,我们评估了12个先进检测器,其中AASIST2达到最优等错误率(EER)13.50%,其余均超过20%。结果表明,这些检测器在真实场景中面临严峻挑战,性能显著下降。此外,我们开展了超过300名参与者的用户研究,结果与12个检测器及多模态大语言模型Qwen2-Audio的表现对比显示,不同检测器与人类在不同欺骗层级下表现出差异化的识别能力,而该大模型完全无法识别深伪语音。我们还公开了检测排行榜,网址为:https://voicewukong.github.io。

原文摘要 · Abstract (English)

With the rapid advancement of technologies like text-to-speech (TTS) and voice conversion (VC), detecting deepfake voices has become increasingly crucial. However, both academia and industry lack a comprehensive and intuitive benchmark for evaluating detectors. Existing datasets are limited in language diversity and lack many manipulations encountered in real-world production environments. To fill this gap, we propose VoiceWukong, a benchmark designed to evaluate the performance of deepfake voice detectors. To build the dataset, we first collected deepfake voices generated by 19 advanced and widely recognized commercial tools and 15 open-source tools. We then created 38 data variants covering six types of manipulations, constructing the evaluation dataset for deepfake voice detection. VoiceWukong thus includes 265,200 English and 148,200 Chinese deepfake voice samples. Using VoiceWukong, we evaluated 12 state-of-the-art detectors. AASIST2 achieved the best equal error rate (EER) of 13.50%, while all others exceeded 20%. Our findings reveal that these detectors face significant challenges in real-world applications, with dramatically declining performance. In addition, we conducted a user study with more than 300 participants. The results are compared with the performance of the 12 detectors and a multimodel large language model (MLLM), i.e., Qwen2-Audio, where different detectors and humans exhibit varying identification capabilities for deepfake voices at different deception levels, while the LALM demonstrates no detection ability at all. Furthermore, we provide a leaderboard for deepfake voice detection, publicly available at {https://voicewukong.github.io}.

语音检测深伪语音基准测试安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。