arXiv:2512.04551cs.SDcs.AI2025-12

通过自适应增强与帧级注意力,提升语音情感识别性能。

Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention

  • 结合能量自适应混合与帧级注意力,生成更丰富的情感特征。
  • 在四个数据集上达到领先效果,显著改善类别不平衡问题。
  • 适合需要高精度情感分析的交互系统开发者使用。

语音情感识别(SER)是人机交互中的关键技术,但因情感复杂性和标注数据稀缺,性能提升困难。为此,本文提出一种多损失学习(MLL)框架,融合基于信噪比的能量自适应混合(EAM)方法与帧级注意力模块(FLAM)。EAM利用信噪比引导的数据增强,生成捕捉细微情感差异的多样化语音样本;FLAM强化多帧情感线索的帧级特征提取。MLL策略整合Kullback-Leibler散度、焦点损失、中心损失与监督对比损失,以优化训练过程、缓解类别不平衡并提升特征可分性。在IEMOCAP、MSP-IMPROV、RAVDESS和SAVEE四个常用数据集上评估,结果表明该方法达到当前最优性能,验证了其有效性和鲁棒性。

原文摘要 · Abstract (English)

Speech emotion recognition (SER) is an important technology in human-computer interaction. However, achieving high performance is challenging due to emotional complexity and scarce annotated data. To tackle these challenges, we propose a multi-loss learning (MLL) framework integrating an energy-adaptive mixup (EAM) method and a frame-level attention module (FLAM). The EAM method leverages SNR-based augmentation to generate diverse speech samples capturing subtle emotional variations. FLAM enhances frame-level feature extraction for multi-frame emotional cues. Our MLL strategy combines Kullback-Leibler divergence, focal, center, and supervised contrastive loss to optimize learning, address class imbalance, and improve feature separability. We evaluate our method on four widely used SER datasets: IEMOCAP, MSP-IMPROV, RAVDESS, and SAVEE. The results demonstrate our method achieves state-of-the-art performance, suggesting its effectiveness and robustness.

语音情感多损失注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。