用标注者分歧信息提升语音情绪识别准确率
Learning from Annotation Uncertainty: Entropy-Aware Curriculum for Speech Emotion Recognition

- 基于标注者投票分布设计软标签训练策略
- 在MSP-Podcast 2.0上降低JSD/KLD达12%-18%
- 适合关注人类感知不确定性的情绪识别研究
语音情绪识别(SER)通常依赖于硬标签,忽略了标注者间的分歧。本文在MSP-Podcast 2.0数据集上,采用WavLM-Base多任务模型进行9类情绪与维度VAD识别,比较硬标签训练与主标注者及主-次标注者合并投票分布作为目标的差异。分布式监督显著提升与人类投票分布的一致性,相比硬标签训练,降低JSD/KLD达12%-18%。分析表明,硬标签部分优势源于将模糊语句归入残差'Other'类别,而分布式监督则将不确定性合理分布在各情绪类别中。熵分层评估显示,高模糊性语句仍具挑战性,但分布式监督更准确捕捉了感知不确定性。结果支持从硬标签转向反映听者分歧的软标签。
原文摘要 · Abstract (English)
Speech emotion recognition (SER) often relies on hard consensus labels that collapse annotator disagreement. We study distribution-based supervision for 9-class SER on MSP-Podcast 2.0 using a WavLM-Base multitask model for categorical emotion and dimensional VAD. Hard-label training is compared with targets from primary and merged primary--secondary annotator vote distributions. Distributional objectives improve alignment with human vote distributions, reducing JSD/KLD relative to hard-label training. Analysis shows that hard supervision partly benefits from assigning ambiguous utterances to the residual Other class, whereas distributional supervision redistributes uncertainty across emotion categories. Entropy-stratified evaluation shows that high-ambiguity utterances remain challenging, but distribution-based supervision better captures perceptual uncertainty. These findings support moving beyond hard labels toward targets that reflect listener disagreement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。