用大模型生成情感标注,解决语音情绪识别中人工标注不足问题
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
- 用音频语言大模型生成合成感知代理标注,补充真实标注
- 在IEMOCAP和MSP-Podcast上提升低模糊度情绪分布可靠性
- 适合关注情绪识别标注瓶颈与多模态数据增强的研究者
语音情绪识别模型通常使用单一类别标签,忽视了人类情绪的固有模糊性。模糊情绪识别通过将情绪表示为概率分布来应对这一问题,但进展受限于从稀疏人工标注中推断出的不可靠真实分布。本文探索大型音频-语言模型(ALMs)是否可通过生成高质量合成标注来缓解标注瓶颈。提出一个框架,利用ALMs生成合成感知代理,以增强人工标注,提升真实分布的可靠性。通过统计分析验证合成代理与人类分布的一致性,并通过微调ALMs评估其影响。此外,为解决类别不平衡并实现无偏评估,提出DiME-Aug:一种分布感知的多模态情绪增强策略。在IEMOCAP和MSP-Podcast上的实验表明,合成标注提升了情绪分布质量,尤其在标注一致性高的低模糊区域效果显著;但在高度模糊、人类分歧大的情绪上收益减弱。本工作首次提供证据表明ALMs可缓解模糊情绪识别中的标注稀缺问题,但也指出需更先进的提示或生成策略应对高度模糊情况。
原文摘要 · Abstract (English)
Speech Emotion Recognition models typically use single categorical labels, overlooking the inherent ambiguity of human emotions. Ambiguous Emotion Recognition addresses this by representing emotions as probability distributions, but progress is limited by unreliable ground-truth distributions inferred from sparse human annotations. This paper explores whether Large Audio-Language Models (ALMs) can mitigate the annotation bottleneck by generating high-quality synthetic annotations. We introduce a framework leveraging ALMs to create Synthetic Perceptual Proxies, augmenting human annotations to improve ground-truth distribution reliability. We validate these proxies through statistical analysis of their alignment with human distributions and evaluate their impact by fine-tuning ALMs with the augmented emotion distributions. Furthermore, to address class imbalance and enable unbiased evaluation, we propose DiME-Aug, a Distribution-aware Multimodal Emotion Augmentation strategy. Experiments on IEMOCAP and MSP-Podcast show that synthetic annotations enhance emotion distribution, especially in low-ambiguity regions where annotation agreement is high. However, benefits diminish for highly ambiguous emotions with greater human disagreement. This work provides the first evidence that ALMs could address annotation scarcity in ambiguous emotion recognition, but highlights the need for more advanced prompting or generation strategies to handle highly ambiguous cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。