arXiv:2506.19887eess.AScs.AI2025-06中稿 · INTERSPEECH 2025被引 3

融合声学与文本多层级特征,提升自然语音情感识别准确率

MATER: Multi-level Acoustic and Textual Emotion Representation for Interpretable Speech Emotion Recognition

  • 构建词、话语、嵌入三级融合框架,捕捉语音细粒度变化与语义内涵
  • 在自然场景下实现41.01%宏平均F1和0.5928平均CCC,valence预测达0.6941
  • 引入不确定性感知集成策略,有效缓解标注不一致问题

本文针对自然语境下的语音情感识别挑战(SERNC),提出多层级声学-文本情感表征(MATER)框架,解决自然语音中个体间与个体内差异带来的复杂性。该框架在词、话语和嵌入三个层级融合声学与文本特征,结合低层词汇与声学线索及高层上下文表示,有效捕捉细微语调变化与语义细节。同时,设计不确定性感知集成策略,降低标注不一致带来的干扰。在两个任务中均位列第四,宏平均F1为41.01%,平均一致性相关系数(CCC)为0.5928;其中在愉悦度(valence)预测上取得第二名,CCC达0.6941。

原文摘要 · Abstract (English)

This paper presents our contributions to the Speech Emotion Recognition in Naturalistic Conditions (SERNC) Challenge, where we address categorical emotion recognition and emotional attribute prediction. To handle the complexities of natural speech, including intra- and inter-subject variability, we propose Multi-level Acoustic-Textual Emotion Representation (MATER), a novel hierarchical framework that integrates acoustic and textual features at the word, utterance, and embedding levels. By fusing low-level lexical and acoustic cues with high-level contextualized representations, MATER effectively captures both fine-grained prosodic variations and semantic nuances. Additionally, we introduce an uncertainty-aware ensemble strategy to mitigate annotator inconsistencies, improving robustness in ambiguous emotional expressions. MATER ranks fourth in both tasks with a Macro-F1 of 41.01% and an average CCC of 0.5928, securing second place in valence prediction with an impressive CCC of 0.6941.

语音情感识别多模态融合自然语境可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。