arXiv:2503.22711cs.SDcs.AI2025-03ICLR被引 2

用情绪评分分布替代共识标签,提升语音情感识别模型性能。

Modeling speech emotion with label variance and analyzing performance across speakers and unseen acoustic conditions

  • 以评分概率密度代替共识标签,更好捕捉情绪不确定性。
  • 新方法在多个基准测试中优于现有文献结果。
  • 跨说话人、性别和声学条件的评估更真实反映模型能力。

自发语音情感数据通常包含感知评分,评分者听录音后打分。这种评分因评分者主观差异带来标签不确定性。现有方法采用共识评分作为真实标签,即选择得票最高的情绪。但共识评分忽略多情绪共存的模糊情况,无法反映评分者的判断不确定性。本文证明,使用情绪评分的概率密度函数作为目标,相比传统共识标签,在基准测试集上表现更优。我们发现,基于显著性驱动的基座模型(FM)表示选择,可训练出当前最优的语音情感识别模型,适用于维度与分类情感识别。对比不同基座模型的表示,发现仅看整体测试集表现会掩盖模型在跨说话人和性别上的泛化能力缺陷。通过多测试集评估与跨性别、说话人分析,能更准确评估模型价值。最后,我们指出标签不确定性与数据偏移对模型评估构成挑战,建议考虑2-3个最佳假设而非单一最优假设。

原文摘要 · Abstract (English)

Spontaneous speech emotion data usually contain perceptual grades where graders assign emotion score after listening to the speech files. Such perceptual grades introduce uncertainty in labels due to grader opinion variation. Grader variation is addressed by using consensus grades as groundtruth, where the emotion with the highest vote is selected. Consensus grades fail to consider ambiguous instances where a speech sample may contain multiple emotions, as captured through grader opinion uncertainty. We demonstrate that using the probability density function of the emotion grades as targets instead of the commonly used consensus grades, provide better performance on benchmark evaluation sets compared to results reported in the literature. We show that a saliency driven foundation model (FM) representation selection helps to train a state-of-the-art speech emotion model for both dimensional and categorical emotion recognition. Comparing representations obtained from different FMs, we observed that focusing on overall test-set performance can be deceiving, as it fails to reveal the models generalization capacity across speakers and gender. We demonstrate that performance evaluation across multiple test-sets and performance analysis across gender and speakers are useful in assessing usefulness of emotion models. Finally, we demonstrate that label uncertainty and data-skew pose a challenge to model evaluation, where instead of using the best hypothesis, it is useful to consider the 2- or 3-best hypotheses.

语音情感标签不确定模型评估基座模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。