arXiv:2505.19693cs.SDcs.AI2025-05被引 2

用球面分区提升语音情感识别准确率

EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification

  • 将情绪维度转为球面坐标,分区域分类辅助回归
  • 在多个数据集上超越基线模型,提升预测一致性
  • 适合关注情感建模与结构化学习的研究者

语音情感识别通过离散标签或连续维度(如唤醒度、效价、支配度,简称VAD)预测说话人情绪状态。本文提出EmoSphere-SER,一种联合模型,将VAD值转换为球面坐标并划分为多个球面区域,引入辅助分类任务预测每个点所属区域,从而指导回归过程。同时,采用动态加权机制和带有多头自注意力的风格池化层,捕捉频谱与时间动态,增强模型表现。联合训练策略强化了结构化学习,提升了预测一致性。实验表明,该方法优于基准模型,验证了框架有效性。

原文摘要 · Abstract (English)

Speech emotion recognition predicts a speaker's emotional state from speech signals using discrete labels or continuous dimensions such as arousal, valence, and dominance (VAD). We propose EmoSphere-SER, a joint model that integrates spherical VAD region classification to guide VAD regression for improved emotion prediction. In our framework, VAD values are transformed into spherical coordinates that are divided into multiple spherical regions, and an auxiliary classification task predicts which spherical region each point belongs to, guiding the regression process. Additionally, we incorporate a dynamic weighting scheme and a style pooling layer with multi-head self-attention to capture spectral and temporal dynamics, further boosting performance. This combined training strategy reinforces structured learning and improves prediction consistency. Experimental results show that our approach exceeds baseline methods, confirming the validity of the proposed framework.

语音情感识别球面表示多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。