arXiv:2510.23506cs.AIcs.HC2025-10AAAI被引 7

让多模态大模型生成与情绪预测一致的解释,提升交互可信度。

Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier

  • 引入情感理由验证器,强制推理过程与目标情绪一致。
  • 在MAFW和DFEW数据集上解释准确率显著提升,预测一致性增强。
  • 无需修改模型或额外标注,适合需可解释情绪识别的场景。

多模态大语言模型(MLLMs)的进步正推动人机交互从表层交流向更细腻、具情感智能的方向演进。实现这一转变的关键在于理解情绪,使系统能够捕捉用户意图背后的微妙线索。同时,对预测情绪提供真实可信的解释,对确保可解释性及建立用户信任至关重要。然而,现有基于MLLM的方法常生成与目标标签不符甚至自相矛盾的情绪解释,这在交互场景中可能导致误解并削弱可靠性。为此,我们提出一种新方法:情感理由验证器(ERV)与解释奖励机制。该方法在不修改模型架构且无需额外视频-描述配对标注的前提下,引导模型在多模态情绪识别中生成与目标情绪显式一致的推理过程。大量实验与人工评估表明,该方法显著提升了解释与预测的一致性及解释情绪准确性,在MAFW和DFEW数据集上表现优异。结果证明,该方法不仅增强了解释与预测的对齐,还使MLLM能提供情感连贯、可信的交互体验,为实现真正类人的交互系统迈出关键一步。

原文摘要 · Abstract (English)

The recent advancement of Multimodal Large Language Models (MLLMs) is transforming human-computer interaction (HCI) from surface-level exchanges into more nuanced and emotionally intelligent communication. To realize this shift, emotion understanding becomes essential allowing systems to capture subtle cues underlying user intent. Furthermore, providing faithful explanations for predicted emotions is crucial to ensure interpretability and build user trust. However, current MLLM-based methods often generate emotion explanations that diverge from the target labels and sometimes even contradict their own predicted emotions. This inconsistency poses a critical risk for misunderstanding and erodes reliability in interactive settings. To address this, we propose a novel approach: the Emotional Rationale Verifier (ERV) and an Explanation Reward. Our method guides the model to produce reasoning that is explicitly consistent with the target emotion during multimodal emotion recognition without modifying the model architecture or requiring additional paired video-description annotations. Our method significantly improves faithful explanation-prediction consistency and explanation emotion accuracy on the MAFW and DFEW datasets. Through extensive experiments and human evaluations, we show that our approach not only enhances alignment between explanation and prediction but also empowers MLLMs to deliver emotionally coherent, trustworthy interactions, marking a key step toward truly human-like HCI systems.

多模态情绪识别可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。