用语音和对话文本联合建模情绪变化,提升对话情绪识别准确率与可解释性。
GatedxLSTM: A Multimodal Affective Computing Approach for Emotion Recognition in Conversations
- 融合说话人与对话伙伴的语音和文本,通过门控机制突出关键语句。
- 在IEMOCAP数据集上四分类准确率达当前开源模型最优水平。
- 结合心理学视角解析情绪演变,适合需要可解释性的对话系统研究者。
情感计算(AC)对推动通用人工智能(AGI)至关重要,情绪识别是其核心组成部分。然而,人类情绪具有内在动态性,不仅受个体表达影响,也受互动关系驱动,单一模态方法难以捕捉其全貌。多模态情绪识别(MER)虽利用多种信号,但传统方法多基于话语级分析,忽略了对话中情绪的动态演化。对话情绪识别(ERC)弥补了这一缺陷,但现有方法仍难以对齐多模态特征并解释情绪变化原因。为此,我们提出GatedxLSTM,一种新型语音-文本多模态对话情绪识别模型,显式考虑说话人及对话伙伴的语音与文本内容,识别驱动情绪转变的关键句子。通过引入对比语言-音频预训练(CLAP)增强跨模态对齐,并采用门控机制突出情绪关键语句,提升模型可解释性与性能。此外,对话情绪解码器(DED)通过建模上下文依赖关系进一步优化预测。在IEMOCAP数据集上的实验表明,GatedxLSTM在四类情绪分类任务中达到当前开源方法的最先进性能,验证了其在实际应用中的有效性,并提供了心理学视角的可解释性分析。
原文摘要 · Abstract (English)
Affective Computing (AC) is essential for advancing Artificial General Intelligence (AGI), with emotion recognition serving as a key component. However, human emotions are inherently dynamic, influenced not only by an individual's expressions but also by interactions with others, and single-modality approaches often fail to capture their full dynamics. Multimodal Emotion Recognition (MER) leverages multiple signals but traditionally relies on utterance-level analysis, overlooking the dynamic nature of emotions in conversations. Emotion Recognition in Conversation (ERC) addresses this limitation, yet existing methods struggle to align multimodal features and explain why emotions evolve within dialogues. To bridge this gap, we propose GatedxLSTM, a novel speech-text multimodal ERC model that explicitly considers voice and transcripts of both the speaker and their conversational partner(s) to identify the most influential sentences driving emotional shifts. By integrating Contrastive Language-Audio Pretraining (CLAP) for improved cross-modal alignment and employing a gating mechanism to emphasise emotionally impactful utterances, GatedxLSTM enhances both interpretability and performance. Additionally, the Dialogical Emotion Decoder (DED) refines emotion predictions by modelling contextual dependencies. Experiments on the IEMOCAP dataset demonstrate that GatedxLSTM achieves state-of-the-art (SOTA) performance among open-source methods in four-class emotion classification. These results validate its effectiveness for ERC applications and provide an interpretability analysis from a psychological perspective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。