arXiv:2508.17878cs.SD2025-08中稿 · interspeech2025被引 2

多任务学习+动态融合,提升语音情感识别准确率

Enhancing Speech Emotion Recognition with Multi-Task Learning and Dynamic Feature Fusion

  • 同时训练情感、性别、说话人识别等4个任务,共享特征表示
  • 在自然场景语音情感数据集上,识别准确率显著提升
  • 自适应调整难样本权重,缓解类别不平衡问题

本研究通过多任务学习微调自监督学习模型,以增强语音情感识别性能。框架同时处理四个相关任务:情感识别、性别识别、说话人验证和自动语音识别。引入创新的共注意力模块,动态捕捉主任务与辅助任务之间的特征交互,实现上下文感知的特征融合。此外,提出样本加权焦点对比损失(SWFC),通过调整困难样本和少数类样本的权重,缓解类别不平衡与语义混淆问题。该方法在自然条件下语音情感识别挑战赛的分类任务上得到验证,表现显著优于基线。

原文摘要 · Abstract (English)

This study investigates fine-tuning self-supervised learn ing (SSL) models using multi-task learning (MTL) to enhance speech emotion recognition (SER). The framework simultane ously handles four related tasks: emotion recognition, gender recognition, speaker verification, and automatic speech recog nition. An innovative co-attention module is introduced to dy namically capture the interactions between features from the primary emotion classification task and auxiliary tasks, en abling context-aware fusion. Moreover, We introduce the Sam ple Weighted Focal Contrastive (SWFC) loss function to ad dress class imbalance and semantic confusion by adjusting sam ple weights for difficult and minority samples. The method is validated on the Categorical Emotion Recognition task of the Speech Emotion Recognition in Naturalistic Conditions Chal lenge, showing significant performance improvements.

语音情感识别多任务学习动态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。