arXiv:2602.23802cs.AIcs.CV2026-02中稿 · CVPR被引 2

让多模态大模型更懂情绪,通过反思式强化学习提升情感推理能力。

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models

  • 引入结构化情感思维,分步引导模型进行可解释的情感推理。
  • 设计反思式情感奖励,基于视觉与文本一致性提升情感连贯性。
  • 在多个视觉情感理解基准上表现更优,增强模型可解释性与情商。

多模态大语言模型(MLLMs)在视觉推理与理解任务中取得显著进展,但仍难以捕捉人类情绪的复杂性与主观性。现有基于监督微调的方法常受限于泛化能力差和可解释性不足,而强化学习方法如组相对策略优化(Group Relative Policy Optimization)未能契合情感认知的内在特性。为此,我们提出面向情感推理的反思式强化学习框架(EMO-R3),旨在提升MLLMs的情感推理能力。具体而言,引入结构化情感思维,引导模型以结构化、可解释的方式逐步进行情感推理;设计反思式情感奖励机制,使模型能基于视觉-文本一致性与情感连贯性重新评估其推理过程。大量实验表明,EMO-R3显著提升了MLLMs的可解释性与情感智能,在多个视觉情感理解基准上均取得优异表现。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to capture the complexity and subjectivity of human emotions. Existing approaches based on supervised fine-tuning often suffer from limited generalization and poor interpretability, while reinforcement learning methods such as Group Relative Policy Optimization fail to align with the intrinsic characteristics of emotional cognition. To address these challenges, we propose Reflective Reinforcement Learning for Emotional Reasoning (EMO-R3), a framework designed to enhance the emotional reasoning ability of MLLMs. Specifically, we introduce Structured Emotional Thinking to guide the model to perform step-by-step emotional reasoning in a structured and interpretable manner, and design a Reflective Emotional Reward that enables the model to re-evaluate its reasoning based on visual-text consistency and emotional coherence. Extensive experiments demonstrate that EMO-R3 significantly improves both the interpretability and emotional intelligence of MLLMs, achieving superior performance across multiple visual emotional understanding benchmarks.

情感推理多模态强化学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。