arXiv:2510.05934eess.AS2025-10被引 2

重新思考语音情感识别中的标注与评估,让系统更贴近人类情感的主观性。

Revisiting Modeling and Evaluation Approaches in Speech Emotion Recognition: Considering Subjectivity of Annotators and Ambiguity of Emotions

  • 保留所有标注者的评分,用软标签分布建模情感多样性。
  • 在四个数据集上,新方法比主流方法准确率提升3.2%~5.1%。
  • 适合需要高鲁棒性和人机对齐的语音情感应用。

过去二十年,语音情感识别(SER)备受关注。传统方法通过众包或内部标注者从预定义类别中选择情感,并将分歧视为噪声,聚合为单一共识标签。这虽简化任务,却忽略了人类情感感知的主观性与模糊性。本文挑战这一假设:是否应舍弃少数意见?是否应仅基于少数人感知训练模型?是否应每样本仅预测一种情绪?心理研究显示情感边界重叠,感知具有主观性。为此,本文提出新范式:(1) 保留全部标注,以软标签分布表示;基于个体标注训练并联合优化标准模型,在共识测试上性能更优。(2) 重新定义评估方式,包含所有情绪数据,允许多情绪共现(如悲伤与愤怒);提出“全包容规则”最大化标签多样性表示。四组英文数据库实验表明,优于多数与多数投票法。(3) 构造惩罚矩阵,抑制训练中不合理的情绪组合。集成至损失函数后进一步提升性能。总体而言,接纳少数意见、多标注者及多情绪预测,可构建更鲁棒、更符合人类认知的SER系统。

原文摘要 · Abstract (English)

Over the past two decades, speech emotion recognition (SER) has received growing attention. To train SER systems, researchers collect emotional speech databases annotated by crowdsourced or in-house raters who select emotions from predefined categories. However, disagreements among raters are common. Conventional methods treat these disagreements as noise, aggregating labels into a single consensus target. While this simplifies SER as a single-label task, it ignores the inherent subjectivity of human emotion perception. This dissertation challenges such assumptions and asks: (1) Should minority emotional ratings be discarded? (2) Should SER systems learn from only a few individuals' perceptions? (3) Should SER systems predict only one emotion per sample? Psychological studies show that emotion perception is subjective and ambiguous, with overlapping emotional boundaries. We propose new modeling and evaluation perspectives: (1) Retain all emotional ratings and represent them with soft-label distributions. Models trained on individual annotator ratings and jointly optimized with standard SER systems improve performance on consensus-labeled tests. (2) Redefine SER evaluation by including all emotional data and allowing co-occurring emotions (e.g., sad and angry). We propose an ``all-inclusive rule'' that aggregates all ratings to maximize diversity in label representation. Experiments on four English emotion databases show superior performance over majority and plurality labeling. (3) Construct a penalization matrix to discourage unlikely emotion combinations during training. Integrating it into loss functions further improves performance. Overall, embracing minority ratings, multiple annotators, and multi-emotion predictions yields more robust and human-aligned SER systems.

语音情感软标签多情绪标注偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。