用后训练提升语音检测模型对新型伪造语音的识别能力
Post-training for Deepfake Speech Detection
- 在通用预训练基础上,通过大规模多语言数据做后训练
- 在超过5.6万小时真实语音上训练,对未见伪造语音表现稳健
- 适合需要高泛化能力的语音安全检测场景
我们提出一种后训练方法,通过弥合通用预训练与领域特定微调之间的差距,提升自监督学习(SSL)模型在深度伪造语音检测中的性能。我们构建了AntiDeepfake系列模型,基于包含超过56,000小时真实语音和18,000小时带各类失真语音的多语言大规模语料库进行训练,覆盖一百多种语言。实验表明,这些后训练模型已具备对未见过的深度伪造语音的强大鲁棒性和泛化能力。当进一步在Deepfake-Eval-2024数据集上微调时,其性能持续超越不采用后训练的现有最先进检测器。模型检查点与源代码已公开。
原文摘要 · Abstract (English)
We introduce a post-training approach that adapts self-supervised learning (SSL) models for deepfake speech detection by bridging the gap between general pre-training and domain-specific fine-tuning. We present AntiDeepfake models, a series of post-trained models developed using a large-scale multilingual speech dataset containing over 56,000 hours of genuine speech and 18,000 hours of speech with various artifacts in over one hundred languages. Experimental results show that the post-trained models already exhibit strong robustness and generalization to unseen deepfake speech. When they are further fine-tuned on the Deepfake-Eval-2024 dataset, these models consistently surpass existing state-of-the-art detectors that do not leverage post-training. Model checkpoints and source code are available online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。