融合Mamba与注意力模型,提升语音情感识别性能
PARROT: Synergizing Mamba and Attention-based SSL Pre-Trained Models via Parallel Branch Hadamard Optimal Transport for Speech Emotion Recognition
- 并行分支结合最优传输与哈达玛积,实现异构预训练模型融合
- 在语音情感识别任务上超越单一模型及同质模型融合方法
- 适合关注多模态预训练模型融合的语音处理研究者
Mamba作为注意力机制的替代架构,推动了基于Mamba的自监督学习(SSL)预训练模型(PTM)在语音与音频处理中的发展。近期研究表明,这些模型在语音情感识别(SER)任务中表现可媲美甚至优于最先进的注意力型PTM。受先前工作启发,即在不同语音任务中融合不同类型的预训练模型具有优势,我们假设:结合Mamba型与注意力型PTM的互补优势,将使SER性能超越仅融合同质注意力型PTM的方法。为此,我们提出新框架PARROT,通过并行分支融合、最优传输与哈达玛积实现异构模型融合。该方法在多项指标上达到当前最优(SOTA)结果,显著优于单个模型、同质模型融合及基线融合技术,验证了异构预训练模型融合在语音情感识别中的潜力。
原文摘要 · Abstract (English)
The emergence of Mamba as an alternative to attention-based architectures has led to the development of Mamba-based self-supervised learning (SSL) pre-trained models (PTMs) for speech and audio processing. Recent studies suggest that these models achieve comparable or superior performance to state-of-the-art (SOTA) attention-based PTMs for speech emotion recognition (SER). Motivated by prior work demonstrating the benefits of PTM fusion across different speech processing tasks, we hypothesize that leveraging the complementary strengths of Mamba-based and attention-based PTMs will enhance SER performance beyond the fusion of homogenous attention-based PTMs. To this end, we introduce a novel framework, PARROT that integrates parallel branch fusion with Optimal Transport and Hadamard Product. Our approach achieves SOTA results against individual PTMs, homogeneous PTMs fusion, and baseline fusion techniques, thus, highlighting the potential of heterogeneous PTM fusion for SER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。