arXiv:2412.05558cs.SDcs.AI2024-12中稿 · 31st International…被引 12

通过跨模态注意力与特征对齐,提升语音情感识别准确率

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition

  • 引入门控交叉注意力捕捉多模态间细微交互
  • 在IEMOCAP和MELD数据集上超越现有最佳方法
  • 适合需要高精度情感分析的多模态系统开发者

语音情感识别(SER)因人类情感的复杂性与多样性仍具挑战。现有方法虽尝试通过多模态学习融合信息,但常忽视跨模态交互的复杂性,导致特征表示不佳。本文提出WavFusion框架,解决有效多模态融合、模态异质性及判别性表征学习三大问题。通过门控交叉注意力机制与多模态同质特征差异学习,WavFusion在两个基准数据集(IEMOCAP和MELD)上表现优于现有最先进方法,验证了捕捉精细跨模态交互与学习判别性表征对精准多模态SER的重要性。

原文摘要 · Abstract (English)

Speech emotion recognition (SER) remains a challenging yet crucial task due to the inherent complexity and diversity of human emotions. To address this problem, researchers attempt to fuse information from other modalities via multimodal learning. However, existing multimodal fusion techniques often overlook the intricacies of cross-modal interactions, resulting in suboptimal feature representations. In this paper, we propose WavFusion, a multimodal speech emotion recognition framework that addresses critical research problems in effective multimodal fusion, heterogeneity among modalities, and discriminative representation learning. By leveraging a gated cross-modal attention mechanism and multimodal homogeneous feature discrepancy learning, WavFusion demonstrates improved performance over existing state-of-the-art methods on benchmark datasets. Our work highlights the importance of capturing nuanced cross-modal interactions and learning discriminative representations for accurate multimodal SER. Experimental results on two benchmark datasets (IEMOCAP and MELD) demonstrate that WavFusion succeeds over the state-of-the-art strategies on emotion recognition.

语音情感识别多模态学习注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。