让模型像人一样推理情感,提升对话情绪识别准确率。
Rationale-Guided Learning for Multimodal Emotion Recognition

- 基于双系统理论拆解情绪推理为直觉、情境、整合三部分
- 在IEMOCAP和MELD上达到当前最优性能
- 能为新样本生成语义正确的原因解释,适合需要可解释性的场景
对话中的多模态情绪识别(MERC)需理解语言与非语言线索间的复杂交互。现有方法多将其视为直接的输入输出映射,忽视人类解读情绪时的因果推理过程。本文提出理性引导学习(RGL),将MERC转化为类人认知推理任务。基于双系统理论,将情绪推理分解为:直觉(即时感知,系统1)、情境(情境分析,系统2)与整合(两者融合)。利用多模态大模型(MLLM)离线生成结构化理由,并作为记忆编码,对齐模型内部表示与人类推理模式。最终模型推理时无需任何MLLM开销。实验表明,RGL在IEMOCAP与MELD基准上达到领先性能。进一步验证显示,模型内部特征可有效检索未见测试样本的语义正确理由,证明其具备理由推理能力。
原文摘要 · Abstract (English)
Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues. However, most existing approaches fundamentally treat this as a direct input-output (multimodal cues-emotion labels) mapping problem, overlooking the causal reasoning that humans use when interpreting emotions. We propose rationale-guided learning (RGL), a novel framework that transforms MERC into a cognitively-inspired reasoning task. Based on dual-process theory, we decompose emotional reasoning into three facets: Intuitive (immediate perception, System 1), Contextual (situational analysis, System 2), and Integrative (synthesis of both). We leverage an MLLM offline to generate structured rationales, which are encoded as memories to guide model training via aligning internal representations with human-like reasoning patterns. Our final model operates without any MLLM overheads at inference time. Experimental results show that RGL achieves state-of-the-art performance on the IEMOCAP and MELD benchmarks. Further, for interpretation, we demonstrate that the model's internal features effectively retrieve semantically correct rationales for unseen test samples, validating its rationale reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。