用WavLM模型检测语音伪造,性能优于现有方法
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
- 以预训练WavLM为前端,结合多种后端技术提取语音特征
- 在开放条件下实现3.42%误报率、0.1927的Cllr得分
- 适合语音安全与反欺诈领域研究人员参考
本文介绍了我们提交至ASVspoof 5挑战赛第1赛道——开放条件下的语音深度伪造检测系统。该任务为独立的真伪语音检测任务。近年来,大规模自监督模型已成为自动语音识别(ASR)及其他语音处理任务的标准工具。为此,我们采用预训练WavLM作为前端模型,并通过不同后端技术对特征表示进行池化。整个框架仅使用挑战赛提供的训练数据进行微调,类似封闭条件设置。此外,我们引入噪声和混响的数据增强策略,分别基于MUSAN噪声数据集和房间脉冲响应(RIR)数据集。还尝试了编解码器增强以提升性能。最终,使用Bosaris工具包进行分数校准与系统融合,以获得更优的Cllr得分。融合系统在测试中达到0.0937 minDCF、3.42% EER、0.1927 Cllr和0.1375 actDCF。
原文摘要 · Abstract (English)
This paper describes our submitted systems to the ASVspoof 5 Challenge Track 1: Speech Deepfake Detection - Open Condition, which consists of a stand-alone speech deepfake (bonafide vs spoof) detection task. Recently, large-scale self-supervised models become a standard in Automatic Speech Recognition (ASR) and other speech processing tasks. Thus, we leverage a pre-trained WavLM as a front-end model and pool its representations with different back-end techniques. The complete framework is fine-tuned using only the trained dataset of the challenge, similar to the close condition. Besides, we adopt data-augmentation by adding noise and reverberation using MUSAN noise and RIR datasets. We also experiment with codec augmentations to increase the performance of our method. Ultimately, we use the Bosaris toolkit for score calibration and system fusion to get better Cllr scores. Our fused system achieves 0.0937 minDCF, 3.42% EER, 0.1927 Cllr, and 0.1375 actDCF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。