arXiv:2409.18558cs.SDeess.AS2024-09被引 9

融合XLS-R与WavLM的混合模型,实现歌声伪造检测新纪录。

XWSB: A Blend System Utilizing XLS-R and WavLM with SLS Classifier detection system for SVDD 2024 Challenge

  • 用XLS-R和WavLM分别提取特征,结合SLS分类器提升识别能力。
  • 在CtrSVDD赛道上达到2.32%的等错误率(EER),性能领先。
  • 适合关注语音伪造检测与多模型融合的研究者参考。

本文介绍了参加2024年歌唱语音深度伪造检测(SVDD)挑战赛所采用的模型结构。该挑战赛为首次设立,旨在应对因非正式语调和不同语速带来的复杂性。本文提出XWSB系统,即XLS-R、WavLM与SLS分类器融合架构,在比赛中取得最佳表现。具体而言,采用在ASVspoof DF数据集上表现最优的XLS-R&SLS结构,并将SLS应用于WavLM形成WavLM&SLS结构。最终通过融合两个模型构建出XWSB系统。实验结果表明,该系统在SVDD挑战赛中展现出优异的识别能力,尤其在CtrSVDD赛道上达到2.32%的等错误率(EER)。代码与数据已公开于https://github.com/QiShanZhang/XWSB_for_SVDD2024。

原文摘要 · Abstract (English)

This paper introduces the model structure used in the SVDD 2024 Challenge. The SVDD 2024 challenge has been introduced this year for the first time. Singing voice deepfake detection (SVDD) which faces complexities due to informal speech intonations and varying speech rates. In this paper, we propose the XWSB system, which achieved SOTA per-formance in the SVDD challenge. XWSB stands for XLS-R, WavLM, and SLS Blend, representing the integration of these technologies for the purpose of SVDD. Specifically, we used the best performing model structure XLS-R&SLS from the ASVspoof DF dataset, and applied SLS to WavLM to form the WavLM&SLS structure. Finally, we integrated two models to form the XWSB system. Experimental results show that our system demonstrates advanced recognition capabilities in the SVDD challenge, specifically achieving an EER of 2.32% in the CtrSVDD track. The code and data can be found at https://github.com/QiShanZhang/XWSB_for_ SVDD2024.

语音伪造深度学习模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。