arXiv:2507.12595eess.AS2025-07

用多语言语音大模型提升真假情感识别效果

Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models

  • 融合多语言语音基础模型,利用跨语言训练优势
  • 在同语种和跨语种场景下均超越现有最佳方法
  • 提出THAMA融合机制,实现模型间协同增效

本文研究情感伪造检测(EFD)。我们假设多语言语音基础模型(SFMs)因在多种语言上预训练,能更好理解音调、语调和强度的差异,从而提升检测能力。通过对比分析当前最先进的SFMs,结果表明多语言SFMs在同语种(域内)和跨语种(域外)评估中均表现更优。为此,我们提出THAMA融合方法,基于相关研究中模型融合可提升性能的发现,利用Tucker分解与Hadamard积的互补性实现有效融合。结合协作式多语言SFMs与THAMA,该方法在域内和域外设置下均达到最优性能,超越单个基础模型、基线融合技术及先前最先进方法。

原文摘要 · Abstract (English)

In this work, we address EmoFake Detection (EFD). We hypothesize that multilingual speech foundation models (SFMs) will be particularly effective for EFD due to their pre-training across diverse languages, enabling a nuanced understanding of variations in pitch, tone, and intensity. To validate this, we conduct a comprehensive comparative analysis of state-of-the-art (SOTA) SFMs. Our results shows the superiority of multilingual SFMs for same language (in-domain) as well as cross-lingual (out-domain) evaluation. To our end, we also propose, THAMA for fusion of foundation models (FMs) motivated by related research where combining FMs have shown improved performance. THAMA leverages the complementary conjunction of tucker decomposition and hadamard product for effective fusion. With THAMA, synergized with cooperative multilingual SFMs achieves topmost performance across in-domain and out-domain settings, outperforming individual FMs, baseline fusion techniques, and prior SOTA methods.

情感识别语音伪造多语言模型模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。