用状态空间模型提升非语言情绪识别效果,融合模型更优。
Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?
- 采用Mamba架构的音频基础模型捕捉情绪内在结构。
- 融合MAFM与AAFMs的RENOM模型达到最新最佳性能。
- 适合关注情感计算与多模态融合的研究者。
本文聚焦于非语言语音情绪识别(NVER)。首次研究基于Mamba的音频基础模型(MAFMs)在NVER中的表现,假设其凭借状态空间建模能更有效地捕捉内在情绪结构,优于依赖注意力机制的音频基础模型(AAFMs),后者可能放大无关模式。MAFMs能提取更稳定、上下文感知的表示,更好区分细微的非语言情绪线索。通过与前沿的AAFMs和MAFMs对比实验,验证了该假设。此外,受语音情绪识别、合成语音检测等研究启发,探索基础模型融合在NVER中的潜力。为此提出RENOM,采用Rényi散度作为新型损失函数实现模型有效对齐,并引入自注意力增强模型内部表示交互。通过异构融合MAFMs与AAFMs,RENOM在个体模型、模型融合及先前最佳方法中均取得最优表现。
原文摘要 · Abstract (English)
In this work, we focus on non-verbal vocal sounds emotion recognition (NVER). We investigate mamba-based audio foundation models (MAFMs) for the first time for NVER and hypothesize that MAFMs will outperform attention-based audio foundation models (AAFMs) for NVER by leveraging its state-space modeling to capture intrinsic emotional structures more effectively. Unlike AAFMs, which may amplify irrelevant patterns due to their attention mechanisms, MAFMs will extract more stable and context-aware representations, enabling better differentiation of subtle non-verbal emotional cues. Our experiments with state-of-the-art (SOTA) AAFMs and MAFMs validates our hypothesis. Further, motivated from related research such as speech emotion recognition, synthetic speech detection, where fusion of foundation models (FMs) have showed improved performance, we also explore fusion of FMs for NVER. To this end, we propose, RENO, that uses renyi-divergence as a novel loss function for effective alignment of the FMs. It also makes use of self-attention for better intra-representation interaction of the FMs. With RENO, through the heterogeneous fusion of MAFMs and AAFMs, we show the topmost performance in comparison to individual FMs, its fusion and also setting SOTA in comparison to previous SOTA work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。