arXiv:2506.10930cs.LG2025-06被引 5

通过多模态与多任务学习,提升自然语音情感识别准确率。

Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction

  • 融合文本嵌入与性别预测,增强模型对情感的感知能力。
  • 在MSP-Podcast数据集上取得挑战赛第一、二名,性能领先。
  • 适合关注真实场景下情感分析的研究者与工程师。

自然语境下的语音情感识别(SER)对语音处理领域构成重大挑战,主要问题包括标注者间意见分歧及数据分布不均。本文提出一个可复现的框架,在语音情感识别自然语境挑战赛(IS25-SER Challenge)-任务2中取得顶级(第一)表现,评估基于MSP-Podcast数据集。系统通过多模态学习、多任务学习及不平衡数据处理策略应对上述挑战。最佳模型通过引入文本嵌入、预测性别,并将‘其他’(O)和‘无共识’(X)样本纳入训练集实现。该系统在挑战赛中斩获第一、第二名,最高性能由简单两系统集成达成。

原文摘要 · Abstract (English)

Speech emotion recognition (SER) in naturalistic conditions presents a significant challenge for the speech processing community. Challenges include disagreement in labeling among annotators and imbalanced data distributions. This paper presents a reproducible framework that achieves superior (top 1) performance in the Emotion Recognition in Naturalistic Conditions Challenge (IS25-SER Challenge) - Task 2, evaluated on the MSP-Podcast dataset. Our system is designed to tackle the aforementioned challenges through multimodal learning, multi-task learning, and imbalanced data handling. Specifically, our best system is trained by adding text embeddings, predicting gender, and including ``Other'' (O) and ``No Agreement'' (X) samples in the training set. Our system's results secured both first and second places in the IS25-SER Challenge, and the top performance was achieved by a simple two-system ensemble.

语音情感识别多模态学习自然语境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。