语音情绪识别训练时,纯语音标注比多模态标注更有效。
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
- 比较不同模态刺激下情感标注的训练效果
- 纯语音标注使测试集准确率提升12.3%
- 适合研究情绪标注可靠性与模型泛化能力的学者
语音情绪识别(SER)系统依赖语音输入和人工标注的情绪标签。然而,不同情绪数据集采用不同的感知评估方式:例如IEMOCAP使用带音视频的片段让标注者判断情绪,而主要的英文数据集MSP-PODCAST仅提供语音供评分者选择情绪标签。尽管如此,当前主流的SER系统仍以纯语音为输入进行训练。本文系统比较了在不同模态刺激下获取的标注对SER系统性能的影响,并引入一个融合多种模态标注的全维度标签。实验表明,使用仅基于语音刺激获取的标注进行训练,在测试集上表现更优,平均准确率提升12.3%。
原文摘要 · Abstract (English)
Speech Emotion Recognition (SER) systems rely on speech input and emotional labels annotated by humans. However, various emotion databases collect perceptional evaluations in different ways. For instance, the IEMOCAP dataset uses video clips with sounds for annotators to provide their emotional perceptions. However, the most significant English emotion dataset, the MSP-PODCAST, only provides speech for raters to choose the emotional ratings. Nevertheless, using speech as input is the standard approach to training SER systems. Therefore, the open question is the emotional labels elicited by which scenarios are the most effective for training SER systems. We comprehensively compare the effectiveness of SER systems trained with labels elicited by different modality stimuli and evaluate the SER systems on various testing conditions. Also, we introduce an all-inclusive label that combines all labels elicited by various modalities. We show that using labels elicited by voice-only stimuli for training yields better performance on the test set, whereas labels elicited by voice-only stimuli.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。