融合语音与音乐大模型,显著提升歌声深度伪造检测效果
Are Music Foundation Models Better at Singing Voice Deepfake Detection? Far-Better Fuse them with Speech Foundation Models
- 用语音大模型提取音高、音调等关键声学特征,优于纯音乐大模型
- 提出FIONA框架,实现双模型同步融合,错误率降至13.74%
- 适合语音安全、音频鉴伪方向的研究者和应用开发者
本研究首次系统比较音乐基础模型(MFMs)与语音基础模型(SFMs)在歌声深度伪造检测(SVDD)中的表现。我们对当前最先进的MFMs(MERT变体和music2vec)及SFMs(通用语音表征学习与说话人识别预训练模型)进行了全面对比。结果表明,说话人识别类SFMs在所有基础模型中表现最佳,因其更有效捕捉了歌声中的音高、音调、强度等特征。为进一步提升性能,我们提出新型融合框架FIONA,通过同步x-vector(说话人识别SFMs)与MERT-v1-330M(MFMs),实现最优检测效果,获得最低等错误率(EER)13.74%,超越所有单一模型及基线融合方法,达到当前最先进水平。
原文摘要 · Abstract (English)
In this study, for the first time, we extensively investigate whether music foundation models (MFMs) or speech foundation models (SFMs) work better for singing voice deepfake detection (SVDD), which has recently attracted attention in the research community. For this, we perform a comprehensive comparative study of state-of-the-art (SOTA) MFMs (MERT variants and music2vec) and SFMs (pre-trained for general speech representation learning as well as speaker recognition). We show that speaker recognition SFM representations perform the best amongst all the foundation models (FMs), and this performance can be attributed to its higher efficacy in capturing the pitch, tone, intensity, etc, characteristics present in singing voices. To our end, we also explore the fusion of FMs for exploiting their complementary behavior for improved SVDD, and we propose a novel framework, FIONA for the same. With FIONA, through the synchronization of x-vector (speaker recognition SFM) and MERT-v1-330M (MFM), we report the best performance with the lowest Equal Error Rate (EER) of 13.74 %, beating all the individual FMs as well as baseline FM fusions and achieving SOTA results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。