arXiv:2506.21277cs.CVcs.CL2025-06被引 51

用强化学习提升多模态模型理解人类意图的全局推理能力

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

  • 通过上下文奖励+逻辑奖励设计,引导模型全面理解多模态信息
  • 在IntentBench上超越现有开源多模态模型,正确率显著提升
  • 适合研究多模态推理、人机交互与情感理解的开发者

随着多模态大语言模型的快速发展,深入理解与解析人类意图已成为关键能力,需依赖细致的推理过程。现有研究中,强化学习(RL)在增强大语言模型(LLM)推理能力方面展现出潜力,但其在多模态数据与格式上的适配仍缺乏探索。本文指出当前多模态推理模型存在两大问题:全局上下文理解不足,以及捷径学习问题。前者导致模型误读多模态上下文而答错;后者表现为模型忽略关键多模态线索,直接回应问题而不整合信息。为此,我们强调模型必须基于对多模态输入全局上下文的清晰理解进行推理,以防止遗漏关键线索并保障推理完整性。为实现准确的上下文理解,我们引入由大语言模型评判的上下文奖励,辅以格式与准确性奖励。同时,利用大语言模型评估逻辑奖励,判断推理过程是否成功融合多模态信息与逻辑方法。此外,我们构建了多模态推理基准测试集IntentBench,用于评估模型对复杂人类意图与情绪的理解能力。实验表明,所提方法在多个多模态基准上优于其他开源模型。

原文摘要 · Abstract (English)

With the rapid evolution of multimodal large language models, the capacity to deeply understand and interpret human intentions has emerged as a critical capability, which demands detailed and thoughtful reasoning. In recent studies, Reinforcement Learning (RL) has demonstrated potential in enhancing the reasoning capabilities of Large Language Models (LLMs). Nonetheless, the challenges associated with adapting RL to multimodal data and formats remain largely unaddressed. In this paper, we identify two issues in existing multimodal reasoning models: insufficient global context understanding and shortcut problems. Insufficient context understanding can happen when a model misinterprets multimodal context, resulting in incorrect answers. The shortcut problem occurs when the model overlooks crucial clues in multimodal inputs, directly addressing the query without considering the multimodal information. To tackle these issues, we emphasize the necessity for the model to reason with a clear understanding of the global context within multimodal inputs. This global context understanding can effectively prevent the model from overlooking key multimodal cues and ensure a thorough reasoning process. To ensure the accurate interpretation of multimodal context information, we implement a context reward judged by a large language model, alongside format and accuracy rewards. Additionally, to improve complex reasoning capability, we employ the LLM to assess the logical reward, determining whether the reasoning process successfully integrates multimodal information with logical methods. We also introduce a reasoning omni-modal benchmark, IntentBench, aimed at evaluating models in understanding complex human intentions and emotions. Our proposed method demonstrates advanced performance across multiple omni-modal benchmarks compared to other open-source omni-modal models.

多模态推理强化学习意图理解大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。