构建自然对话中的情感标签数据集,提升语音情感识别真实性和可解释性。
Switchboard-Affect: Emotion Perception Labels from Conversational Speech
- 基于真实对话语料,训练众包标注者识别10类情绪和3维情感属性。
- 模型在愤怒情绪上表现最差,表明现有方法对自然情绪泛化能力不足。
- 适合研究真实场景下语音情感识别的学者与开发者使用。
理解语音情感数据集构建与标注的细微差别,对于评估语音情感识别(SER)模型在真实应用中的潜力至关重要。大多数训练与评测数据集包含表演或伪表演语音(如播客语音),其中情感表达可能被夸大或有意调整。此外,基于大众感知的标注数据集常缺乏标注者指南的透明度。这些因素使得难以准确评估模型性能并定位改进方向。为弥补这一空白,我们选定Switchboard语料库作为自然对话语音的优质来源,并训练众包标注者对其中语音进行分类情绪(愤怒、轻蔑、厌恶、恐惧、悲伤、惊讶、快乐、温柔、平静、中性)及维度属性(激活度、效价、支配度)标注。该标注集称为Switchboard-Affect(SWB-Affect)。本文详细阐述了标注方法,包括提供给标注者的定义,以及影响其感知的词汇和副语言线索分析。同时,我们评估了当前先进的SER模型,发现不同情绪类别间性能差异显著,尤其在愤怒情绪上泛化能力最差。这些结果凸显了使用捕捉自然情感变化的数据集进行评估的重要性。我们已公开发布SWB-Affect的标注数据,以促进该领域的进一步研究。
原文摘要 · Abstract (English)
Understanding the nuances of speech emotion dataset curation and labeling is essential for assessing speech emotion recognition (SER) model potential in real-world applications. Most training and evaluation datasets contain acted or pseudo-acted speech (e.g., podcast speech) in which emotion expressions may be exaggerated or otherwise intentionally modified. Furthermore, datasets labeled based on crowd perception often lack transparency regarding the guidelines given to annotators. These factors make it difficult to understand model performance and pinpoint necessary areas for improvement. To address this gap, we identified the Switchboard corpus as a promising source of naturalistic conversational speech, and we trained a crowd to label the dataset for categorical emotions (anger, contempt, disgust, fear, sadness, surprise, happiness, tenderness, calmness, and neutral) and dimensional attributes (activation, valence, and dominance). We refer to this label set as Switchboard-Affect (SWB-Affect). In this work, we present our approach in detail, including the definitions provided to annotators and an analysis of the lexical and paralinguistic cues that may have played a role in their perception. In addition, we evaluate state-of-the-art SER models, and we find variable performance across the emotion categories with especially poor generalization for anger. These findings underscore the importance of evaluation with datasets that capture natural affective variations in speech. We release the labels for SWB-Affect to enable further analysis in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。