提出新模型提升多模态情绪识别准确率
GIA-MIC: Multimodal Emotion Recognition with Gated Interactive Attention and Modality-Invariant Learning Constraints
- 设计门控交互注意力,自适应提取各模态特征
- 引入模态不变生成器,对齐跨模态相似性
- 在IEMOCAP上达到80.7%加权准确率,适合人机交互
多模态情绪识别(MER)从视觉、语音和文本等多源数据中提取情绪,在人机交互中具有重要作用。当前主流方法基于注意力融合,虽表现优异,但仍面临两大挑战:如何有效提取模态特异性特征,以及在模态异质性导致的分布差异下捕捉跨模态相似性。为此,本文提出门控交互注意力机制,通过成对交互自适应提取模态特异性特征并增强情感信息;同时引入模态不变生成器,学习模态不变表示,并通过对齐跨模态相似性约束领域偏移。在IEMOCAP数据集上的实验表明,该方法优于现有先进模型,达到加权准确率80.7%与无加权准确率81.3%。
原文摘要 · Abstract (English)
Multimodal emotion recognition (MER) extracts emotions from multimodal data, including visual, speech, and text inputs, playing a key role in human-computer interaction. Attention-based fusion methods dominate MER research, achieving strong classification performance. However, two key challenges remain: effectively extracting modality-specific features and capturing cross-modal similarities despite distribution differences caused by modality heterogeneity. To address these, we propose a gated interactive attention mechanism to adaptively extract modality-specific features while enhancing emotional information through pairwise interactions. Additionally, we introduce a modality-invariant generator to learn modality-invariant representations and constrain domain shifts by aligning cross-modal similarities. Experiments on IEMOCAP demonstrate that our method outperforms state-of-the-art MER approaches, achieving WA 80.7% and UA 81.3%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。