让AI先关注关键感官信息,提升情感理解的可靠性。
Learning What to Attend First: Modality-Importance-Guided Reasoning for Reliable Multimodal Emotion Understanding
- 根据情绪主导模态重排推理顺序,避免被无关信息误导
- 在DFEW数据集上,错误解释率从18.10%降至7.37%
- 适合需要可解释、可靠多模态情感分析的应用场景
本文提出模态重要性引导推理(MIGR)框架,旨在提升多模态大模型在情感理解中的推理可靠性。现有方法常出现推理漂移:模型逐渐依赖自身生成文本而非多模态证据,且解释受视觉路径过度影响。为此,我们引入模态重要性(MI)机制,识别对目标情绪起主导作用的模态。MIGR通过两阶段框架——模态对齐的监督微调与模态感知的奖励优化——确保推理从情绪主导模态开始,生成情感扎根、因果相关且连贯的解释。在DFEW基准上的实验表明,正确预测但解释不一致的情况从18.10%显著降低至7.37%,验证了从主导模态启动推理的有效性。
原文摘要 · Abstract (English)
In this paper, we present Modality-Importance-Guided Reasoning (MIGR), a framework designed to improve the reliability of reasoning-based multimodal emotion understanding in multimodal large language models. Although existing methods have advanced emotion understanding, they often suffer from reasoning drift: models gradually rely on their own generated text instead of multimodal evidence, and their explanations are overly shaped by visually initiated reasoning paths. To address these issues, we introduce Modality Importance (MI), a simple yet effective mechanism for identifying the emotion-dominant modality. Using MI, MIGR reorganizes reasoning sequences so that explanations begin from the modality most critical to the target emotion, preventing early reasoning from being misled by less informative cues. Our two-stage framework-comprising modality-aligned supervised fine-tuning and modality-aware reward optimization-encourages models to generate emotionally grounded, causally relevant, and coherence-preserving explanations. Experimental results on the DFEW benchmark show that MIGR substantially improves reasoning reliability, decreasing instances of correct predictions accompanied by emotionally inconsistent explanations from 18.10% to 7.37%. These results confirm the benefit of initiating reasoning from the emotion-dominant modality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。