构建超大规模语音伪造检测数据集,提升模型真实场景泛化能力
AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds

- 构建含300万片段的多模型合成语音数据集
- 新数据集上现有模型泛化性能差,跨域检测准确率达EER 1.87%
- 提出课程学习法缓解多源伪造导致的负迁移问题
语音合成系统现已能生成高度逼真的声音,带来严峻的真实性挑战。尽管深度伪造检测模型取得进展,其真实世界效果常受训练与测试数据分布变化影响,源于人类语音的复杂性及合成技术的快速演进。现有数据集存在真实语音多样性不足、近期合成系统覆盖有限、伪造来源混杂等问题,阻碍系统评估与开放世界模型训练。为此,我们提出AUDETER(AUdio DEepfake TEst Range)数据集,包含超过4,500小时由11个最新文本转语音模型和10种声码器生成的合成音频,总计300万段。我们发现,多数现有检测器采用二元监督训练,在训练数据包含高度多样伪造模式时,易引发跨源负迁移,影响整体泛化。作为补充,我们提出一种基于课程学习的有效方法以缓解该问题。大量实验表明,现有检测模型难以泛化至AUDETER中的新型伪造语音与真实语音,而基于XLR的检测器在该数据集上训练后,跨域表现优异,在In-the-Wild基准上达到EER 1.87%。AUDETER已开源于GitHub。
原文摘要 · Abstract (English)
Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, their real-world effectiveness is often undermined by evolving distribution shifts between training and test data, driven by the complexity of human speech and the rapid evolution of synthesis systems. Existing datasets suffer from limited real speech diversity, insufficient coverage of recent synthesis systems, and heterogeneous mixtures of deepfake sources, which hinder systematic evaluation and open-world model training. To address these issues, we introduce AUDETER (AUdio DEepfake TEst Range), a large-scale and highly diverse deepfake audio dataset comprising over 4,500 hours of synthetic audio generated by 11 recent TTS models and 10 vocoders, totalling 3 million clips. We further observe that most existing detectors default to binary supervised training, which can induce negative transfer across synthesis sources when the training data contains highly diverse deepfake patterns, impacting overall generalisation. As a complementary contribution, we propose an effective curriculum-learning-based approach to mitigate this effect. Extensive experiments show that existing detection models struggle to generalise to novel deepfakes and human speech in AUDETER, whereas XLR-based detectors trained on AUDETER achieve strong cross-domain performance across multiple benchmarks, achieving an EER of 1.87% on In-the-Wild. AUDETER is available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。