arXiv:2503.06405cs.SDcs.AI2025-03被引 1

解决语音情感识别中音频与文本模态的差异问题,提升跨模态融合效果。

Heterogeneous bimodal attention fusion for speech emotion recognition

  • 引入动态双模注意力与门控机制,实现多层级跨模态交互。
  • 在MELD和IEMOCAP数据集上准确率超越现有最佳方法。
  • 适合关注跨模态融合与情感计算的研究者阅读。

对话中的多模态情感识别因不同模态间复杂且互补的交互而具有挑战性。音频与文本线索对理解人类情感尤为重要。现有研究多聚焦于同一表征层次上的音视频与文本交互,却常忽视低层音频表示与高层文本表示之间的异构模态差距。为此,本文提出一种新型框架——异构双模注意力融合(HBAF),用于对话情感识别中的多层次多模态交互。该方法包含三个核心模块:单模态表征模块、多模态融合模块和跨模态对比学习模块。单模态表征模块将上下文信息融入低层音频表示,以弥合异构模态差距,促进更有效的融合。多模态融合模块采用动态双模注意力与动态门控机制,过滤错误的跨模态关联,充分挖掘模内与模间交互。跨模态对比学习模块捕捉音频与文本模态间的复杂绝对与相对关系。在MELD与IEMOCAP数据集上的实验表明,所提HBAF方法优于现有最先进基线模型。

原文摘要 · Abstract (English)

Multi-modal emotion recognition in conversations is a challenging problem due to the complex and complementary interactions between different modalities. Audio and textual cues are particularly important for understanding emotions from a human perspective. Most existing studies focus on exploring interactions between audio and text modalities at the same representation level. However, a critical issue is often overlooked: the heterogeneous modality gap between low-level audio representations and high-level text representations. To address this problem, we propose a novel framework called Heterogeneous Bimodal Attention Fusion (HBAF) for multi-level multi-modal interaction in conversational emotion recognition. The proposed method comprises three key modules: the uni-modal representation module, the multi-modal fusion module, and the inter-modal contrastive learning module. The uni-modal representation module incorporates contextual content into low-level audio representations to bridge the heterogeneous multi-modal gap, enabling more effective fusion. The multi-modal fusion module uses dynamic bimodal attention and a dynamic gating mechanism to filter incorrect cross-modal relationships and fully exploit both intra-modal and inter-modal interactions. Finally, the inter-modal contrastive learning module captures complex absolute and relative interactions between audio and text modalities. Experiments on the MELD and IEMOCAP datasets demonstrate that the proposed HBAF method outperforms existing state-of-the-art baselines.

情感识别跨模态融合注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。