用多粒度语义增强语音情感识别,提升细粒度情绪预测效果
MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition
- 引入局部强调、全局和扩展三层次语义,融合到声学特征中
- 在MSP-Podcast和IEMOCAP上实现更优的维度化情绪预测性能
- 适合关注情绪细微变化与文本深层语义的语音分析研究者
连续维度语音情感识别通过效价、唤醒度和支配感捕捉情感变化,比分类方法提供更细粒度表征。然而多数多模态方法仅依赖全局文本,存在两大局限:(1) 所有词汇同等处理,忽视句子不同部分的强调会改变情感含义;(2) 仅表达表面词法内容,缺乏高层解释性线索。为此,我们提出MSF-SER(多粒度语义融合语音情感识别),通过局部强调语义(LES)、全局语义(GS)和扩展语义(ES)三层互补文本语义,增强声学特征。采用模内门控融合与跨模态FiLM调制的轻量级专家混合(FM-MOE)进行融合。在MSP-Podcast和IEMOCAP数据集上的实验表明,MSF-SER持续提升维度化预测表现,验证了丰富语义融合在语音情感识别中的有效性。
原文摘要 · Abstract (English)
Continuous dimensional speech emotion recognition captures affective variation along valence, arousal, and dominance, providing finer-grained representations than categorical approaches. Yet most multimodal methods rely solely on global transcripts, leading to two limitations: (1) all words are treated equally, overlooking that emphasis on different parts of a sentence can shift emotional meaning; (2) only surface lexical content is represented, lacking higher-level interpretive cues. To overcome these issues, we propose MSF-SER (Multi-granularity Semantic Fusion for Speech Emotion Recognition), which augments acoustic features with three complementary levels of textual semantics--Local Emphasized Semantics (LES), Global Semantics (GS), and Extended Semantics (ES). These are integrated via an intra-modal gated fusion and a cross-modal FiLM-modulated lightweight Mixture-of-Experts (FM-MOE). Experiments on MSP-Podcast and IEMOCAP show that MSF-SER consistently improves dimensional prediction, demonstrating the effectiveness of enriched semantic fusion for SER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。