arXiv:2506.09556cs.CL2025-06被引 4

MEDUSA通过多阶段融合训练,提升自然语境下语音情感识别准确率。

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions

  • 四阶段框架融合声学与语言特征,用深度交叉模态注意力增强表达
  • 在自然场景挑战赛中排名第一,有效缓解类别不平衡与情感模糊问题
  • 适合需要高鲁棒性情感分析的智能客服、心理健康评估等场景

语音情感识别(SER)因人类情绪主观性强且在自然情境中分布不均而极具挑战。本文提出MEDUSA,一种四阶段训练的多模态深度融合框架,有效应对类别不平衡与情感模糊问题。前两阶段训练基于DeepSER的分类器集成,DeepSER是预训练自监督声学与语言表征的新型深度跨模态变换融合机制;引入流形混合增强正则化。后两阶段优化可学习的元分类器,融合集成预测结果。训练过程结合人工标注评分作为软标签,辅以均衡采样和多任务学习。MEDUSA在Interspeech 2025:自然情景语音情感识别挑战赛任务1中排名首位。

原文摘要 · Abstract (English)

SER is a challenging task due to the subjective nature of human emotions and their uneven representation under naturalistic conditions. We propose MEDUSA, a multimodal framework with a four-stage training pipeline, which effectively handles class imbalance and emotion ambiguity. The first two stages train an ensemble of classifiers that utilize DeepSER, a novel extension of a deep cross-modal transformer fusion mechanism from pretrained self-supervised acoustic and linguistic representations. Manifold MixUp is employed for further regularization. The last two stages optimize a trainable meta-classifier that combines the ensemble predictions. Our training approach incorporates human annotation scores as soft targets, coupled with balanced data sampling and multitask learning. MEDUSA ranked 1st in Task 1: Categorical Emotion Recognition in the Interspeech 2025: Speech Emotion Recognition in Naturalistic Conditions Challenge.

语音情感识别多模态融合深度学习自然场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。