arXiv:2504.21277cs.AI2025-04综述被引 36

用强化学习提升多模态大模型的推理能力,系统梳理方法与挑战

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models

  • 分值函数无关与基于值函数两类强化学习方法
  • 通过优化推理路径和对齐跨模态信息增强推理能力
  • 适合研究多模态推理、强化学习融合的学者参考

将强化学习(RL)应用于提升多模态大语言模型(MLLMs)的推理能力已成为快速发展的研究领域。尽管MLLMs扩展了大语言模型(LLMs)处理视觉、音频、视频等多模态输入的能力,但实现鲁棒的跨模态推理仍具挑战。本文系统综述了基于RL的MLLM推理最新进展,涵盖关键算法设计、奖励机制创新及实际应用。重点分析两类主流RL范式:值函数无关与值函数依赖方法,探讨其如何通过优化推理轨迹和对齐多模态信息来增强推理能力。此外,全面梳理基准数据集、评估协议及当前局限,提出未来方向以应对稀疏奖励、低效跨模态推理和真实部署约束等问题。旨在为基于强化学习的多模态推理提供结构化、全面的指引。

原文摘要 · Abstract (English)

The application of reinforcement learning (RL) to enhance the reasoning capabilities of Multimodal Large Language Models (MLLMs) constitutes a rapidly advancing research area. While MLLMs extend Large Language Models (LLMs) to handle diverse modalities such as vision, audio, and video, enabling robust reasoning across multimodal inputs remains challenging. This paper provides a systematic review of recent advances in RL-based reasoning for MLLMs, covering key algorithmic designs, reward mechanism innovations, and practical applications. We highlight two main RL paradigms, value-model-free and value-model-based methods, and analyze how RL enhances reasoning abilities by optimizing reasoning trajectories and aligning multimodal information. Additionally, we provide an extensive overview of benchmark datasets, evaluation protocols, and current limitations, and propose future research directions to address challenges such as sparse rewards, inefficient cross-modal reasoning, and real-world deployment constraints. Our goal is to provide a comprehensive and structured guide to RL-based multimodal reasoning.

多模态强化学习大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。