区分描述性与表达性语义,提升语音情感识别准确性。
Semantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles
- 区分话语内容与情绪表达两种语义类型
- 描述性语义匹配意图情绪,表达性语义对应真实感受
- 为更智能的人机交互系统提供理论支持
语音情感识别(SER)对提升人机交互至关重要,但其准确性受限于语音中情感细微差异的复杂性。本研究区分了描述性语义(反映话语上下文内容)与表达性语义(反映说话者情绪状态)。在观看情绪化电影片段后,参与者录制了描述体验的音频片段,并标注了每段的意图情绪标签、自评情绪反应及愉悦度/唤醒度评分。实验表明,描述性语义与意图情绪一致,而表达性语义与诱发情绪相关。研究结果为面向人-智能体交互的SER应用提供了依据,推动更具情境感知能力的AI系统发展。
原文摘要 · Abstract (English)
Speech Emotion Recognition (SER) is essential for improving human-computer interaction, yet its accuracy remains constrained by the complexity of emotional nuances in speech. In this study, we distinguish between descriptive semantics, which represents the contextual content of speech, and expressive semantics, which reflects the speaker's emotional state. After watching emotionally charged movie segments, we recorded audio clips of participants describing their experiences, along with the intended emotion tags for each clip, participants' self-rated emotional responses, and their valence/arousal scores. Through experiments, we show that descriptive semantics align with intended emotions, while expressive semantics correlate with evoked emotions. Our findings inform SER applications in human-AI interaction and pave the way for more context-aware AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。