arXiv:2503.05858cs.SDcs.AI2025-03

通过三模块注意力融合,提升语音情感识别的多模态交互能力。

Bimodal Connection Attention Fusion for Speech Emotion Recognition

  • 设计交互连接网络,建模音视频与文本间的跨模态关系。
  • 在MELD和IEMOCAP数据集上准确率超越现有最优方法。
  • 适合关注多模态情感分析与跨模态融合的研究者。

多模态情感识别因难以提取细微情感差异特征而面临挑战。理解多模态间交互与关联是构建高效双模态语音情感识别系统的关键。本文提出双模态连接注意力融合(BCAF)方法,包含三个核心模块:交互连接网络、双模态注意力网络和相关性注意力网络。交互连接网络采用编码器-解码器结构,建模音频与文本间的模态关联,同时保留模态特异性特征。双模态注意力网络增强语义互补性,并挖掘模态内与模态间交互。相关性注意力网络降低跨模态噪声,捕捉音频与文本间的相关性。在MELD和IEMOCAP数据集上的实验表明,所提BCAF方法优于现有最先进基线模型。

原文摘要 · Abstract (English)

Multi-modal emotion recognition is challenging due to the difficulty of extracting features that capture subtle emotional differences. Understanding multi-modal interactions and connections is key to building effective bimodal speech emotion recognition systems. In this work, we propose Bimodal Connection Attention Fusion (BCAF) method, which includes three main modules: the interactive connection network, the bimodal attention network, and the correlative attention network. The interactive connection network uses an encoder-decoder architecture to model modality connections between audio and text while leveraging modality-specific features. The bimodal attention network enhances semantic complementation and exploits intra- and inter-modal interactions. The correlative attention network reduces cross-modal noise and captures correlations between audio and text. Experiments on the MELD and IEMOCAP datasets demonstrate that the proposed BCAF method outperforms existing state-of-the-art baselines.

情感识别多模态注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。