arXiv:2505.20007eess.AScs.SD2025-05中稿 · INTERSPEECH 2025被引 6

通过跨模态对齐与平衡堆叠,提升自然语境下语音情感识别准确率

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model

  • 用跨模态注意力融合多模态特征,增强情感表征
  • 加权交叉熵+软边界损失,缓解8类情绪数据不平衡问题
  • 12模型集成的平衡堆叠结构,显著提升泛化性能

情感在人类交互中起着基础作用,因此能够识别语音情感的系统在人机交互中至关重要。语音情感识别(SER)是一项具有挑战性的任务,尤其是在自然语音场景下且各类情绪数据不平衡时。本文介绍了我们在2025年自然情境语音情感识别挑战赛中的解决方案。所提架构利用跨模态信息,通过跨模态注意力机制融合不同模态的表示。为应对类别不平衡,采用两种训练策略:(i) 加权交叉熵损失(WCE);(ii) WCE结合额外的中性表达软边界损失与平衡机制。共训练了12个多模态模型,并使用平衡堆叠模型进行集成。所提系统在8类语音情感识别任务上达到0.4094的宏平均F1值和0.4128的准确率。

原文摘要 · Abstract (English)

Emotion plays a fundamental role in human interaction, and therefore systems capable of identifying emotions in speech are crucial in the context of human-computer interaction. Speech emotion recognition (SER) is a challenging problem, particularly in natural speech and when the available data is imbalanced across emotions. This paper presents our proposed system in the context of the 2025 Speech Emotion Recognition in Naturalistic Conditions Challenge. Our proposed architecture leverages cross-modality, utilizing cross-modal attention to fuse representations from different modalities. To address class imbalance, we employed two training designs: (i) weighted crossentropy loss (WCE); and (ii) WCE with an additional neutralexpressive soft margin loss and balancing. We trained a total of 12 multimodal models, which were ensembled using a balanced stacking model. Our proposed system achieves a MacroF1 score of 0.4094 and an accuracy of 0.4128 on 8-class speech emotion recognition.

语音识别情感分析跨模态模型集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。