arXiv:2506.03364eess.AScs.MM2025-06中稿 · INTERSPEECH 2025被引 4

用多模态模型识别歌声伪造来源,准确率显著提升。

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models

  • 利用多模态基础模型捕捉声音细微特征
  • 融合模型后达到当前最佳识别准确率
  • 适合语音安全与数字内容可信性研究者

本文提出歌唱语音深度伪造源归属(SVDSA)任务。我们假设多模态基础模型(如 ImageBind、LanguageBind)因具备跨模态预训练能力,能更有效捕捉每种歌声伪造源的独特音色、调音偏差或合成痕迹等细微特征。实验验证了该假设:相较于语音或音乐基础模型,多模态基础模型在 SVDSA 上表现最优。此外,受相关研究启发,我们探索模型融合方法,提出新框架 COFFE,采用 Chernoff 距离作为损失函数实现高效融合。通过多模态基础模型的协同作用,COFFE 在所有单个模型及基线融合方法中取得最高性能。

原文摘要 · Abstract (English)

In this work, we introduce the task of singing voice deepfake source attribution (SVDSA). We hypothesize that multimodal foundation models (MMFMs) such as ImageBind, LanguageBind will be most effective for SVDSA as they are better equipped for capturing subtle source-specific characteristics-such as unique timbre, pitch manipulation, or synthesis artifacts of each singing voice deepfake source due to their cross-modality pre-training. Our experiments with MMFMs, speech foundation models and music foundation models verify the hypothesis that MMFMs are the most effective for SVDSA. Furthermore, inspired from related research, we also explore fusion of foundation models (FMs) for improved SVDSA. To this end, we propose a novel framework, COFFE which employs Chernoff Distance as novel loss function for effective fusion of FMs. Through COFFE with the symphony of MMFMs, we attain the topmost performance in comparison to all the individual FMs and baseline fusion methods.

语音伪造多模态源归属深度伪造检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。