arXiv:2507.21463cs.SDeess.AS2025-07ACL被引 21

构建超大规模多语言语音伪造数据集,助力对抗新型深度伪造音频

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

  • 涵盖40种语音生成工具,覆盖文本转语音、语音转换等前沿技术
  • 含超300万样本、超3000小时音频,支持46种语言,规模行业领先
  • 提供检测模型基线与影响因素分析,适合安全研究者和防御开发者

随着语音生成技术发展,深度伪造音频的滥用风险日益突出,亟需强大的检测系统。然而现有语音伪造数据集在规模和多样性上存在局限,难以训练泛化能力强的检测模型。为此,我们提出SpeechFake,一个专为语音伪造检测设计的大规模数据集,包含超过300万条深度伪造样本,总计超过3000小时音频,由40种不同语音合成工具生成。数据集覆盖文本转语音、语音转换、神经声码器等多种生成技术,并融入最新前沿方法,支持46种语言。本文详细介绍了数据集的构建过程、组成与统计信息,通过在SpeechFake上训练检测模型,展示了其在自身测试集及多种未见测试集上的优异表现。同时,我们系统探究了生成方法、语言多样性与说话人差异对检测性能的影响。我们相信,SpeechFake将成为推动语音伪造检测研究的重要资源。

原文摘要 · Abstract (English)

As speech generation technology advances, the risk of misuse through deepfake audio has become a pressing concern, which underscores the critical need for robust detection systems. However, many existing speech deepfake datasets are limited in scale and diversity, making it challenging to train models that can generalize well to unseen deepfakes. To address these gaps, we introduce SpeechFake, a large-scale dataset designed specifically for speech deepfake detection. SpeechFake includes over 3 million deepfake samples, totaling more than 3,000 hours of audio, generated using 40 different speech synthesis tools. The dataset encompasses a wide range of generation techniques, including text-to-speech, voice conversion, and neural vocoder, incorporating the latest cutting-edge methods. It also provides multilingual support, spanning 46 languages. In this paper, we offer a detailed overview of the dataset's creation, composition, and statistics. We also present baseline results by training detection models on SpeechFake, demonstrating strong performance on both its own test sets and various unseen test sets. Additionally, we conduct experiments to rigorously explore how generation methods, language diversity, and speaker variation affect detection performance. We believe SpeechFake will be a valuable resource for advancing speech deepfake detection and developing more robust models for evolving generation techniques.

语音伪造数据集深度伪造多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。