用Mamba提升多模态情感识别,更好对齐与捕捉长对话情绪变化。
Mamba-Enhanced Text-Audio-Video Alignment Network for Emotion Recognition in Conversations
- 引入Mamba结构,高效对齐文本、音频、视频模态特征。
- 在MELD和IEMOCAP上显著超越现有方法,准确率大幅领先。
- 适合研究长序列多模态情感分析的学者与开发者。
对话中的情感识别(ERC)是多模态交互研究的重要方向,旨在准确识别和分类说话者在对话中表达的情绪。传统方法主要依赖单一模态线索——如文本、音频或视觉数据——导致效果受限。这些方法面临两大挑战:1)多模态信息的一致性;在融合不同模态前,需确保各来源数据对齐且一致。2)上下文信息的捕捉;有效融合多模态特征需要深入理解随时间演变的情感基调,尤其在长对话中情绪可能持续变化。为此,本文提出一种新型的基于Mamba的文本-音频-视频对齐网络(MaTAV),兼具对齐单模态特征以保证跨模态一致性、处理长输入序列以更好地捕捉上下文多模态信息的优势。在MELD和IEMOCAP数据集上的大量实验表明,MaTAV在ERC任务上显著优于现有最先进方法,性能提升明显。
原文摘要 · Abstract (English)
Emotion Recognition in Conversations (ERCs) is a vital area within multimodal interaction research, dedicated to accurately identifying and classifying the emotions expressed by speakers throughout a conversation. Traditional ERC approaches predominantly rely on unimodal cues\-such as text, audio, or visual data\-leading to limitations in their effectiveness. These methods encounter two significant challenges: 1) Consistency in multimodal information. Before integrating various modalities, it is crucial to ensure that the data from different sources is aligned and coherent. 2) Contextual information capture. Successfully fusing multimodal features requires a keen understanding of the evolving emotional tone, especially in lengthy dialogues where emotions may shift and develop over time. To address these limitations, we propose a novel Mamba-enhanced Text-Audio-Video alignment network (MaTAV) for the ERC task. MaTAV is with the advantages of aligning unimodal features to ensure consistency across different modalities and handling long input sequences to better capture contextual multimodal information. The extensive experiments on the MELD and IEMOCAP datasets demonstrate that MaTAV significantly outperforms existing state-of-the-art methods on the ERC task with a big margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。