arXiv:2505.18536cs.CLcs.AI2025-05被引 16

强化微调让多模态大模型更会推理

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

  • 用强化微调提升多模态模型的逻辑推理能力
  • 覆盖多种模态、任务和领域,效果显著
  • 适合关注AI推理能力进化的研究者

站在2025年这一通往通用人工智能(AGI)的关键节点上,强化微调(RFT)在提升大语言模型(LLM)推理能力方面展现出巨大潜力,催生了OpenAI-o1、DeepSeek-R1等前沿模型。如今,将RFT高效应用于多模态大语言模型(MLLMs)以增强其推理能力,已成为学界广泛关注的课题。本文主张:强化微调是驱动多模态大模型推理能力的核心动力。首先,我们系统梳理了该领域研究者应掌握的基础背景知识;其次,从五个维度总结了RFT在提升MLLM推理能力上的进展:多样化的模态、多样化任务与领域、更优的训练算法、丰富的评测基准以及蓬勃发展的工程框架;最后,提出五个未来值得探索的研究方向。我们希望本文能为当前迈向AGI的关键阶段提供有益洞见。相关工作综述详见https://github.com/Sun-Haoyuan23/Awesome-RL-based-Reasoning-MLLMs。

原文摘要 · Abstract (English)

Standing in 2025, at a critical juncture in the pursuit of Artificial General Intelligence (AGI), reinforcement fine-tuning (RFT) has demonstrated significant potential in enhancing the reasoning capability of large language models (LLMs) and has led to the development of cutting-edge AI models such as OpenAI-o1 and DeepSeek-R1. Moreover, the efficient application of RFT to enhance the reasoning capability of multimodal large language models (MLLMs) has attracted widespread attention from the community. In this position paper, we argue that reinforcement fine-tuning powers the reasoning capability of multimodal large language models. To begin with, we provide a detailed introduction to the fundamental background knowledge that researchers interested in this field should be familiar with. Furthermore, we meticulously summarize the improvements of RFT in powering reasoning capability of MLLMs into five key points: diverse modalities, diverse tasks and domains, better training algorithms, abundant benchmarks and thriving engineering frameworks. Finally, we propose five promising directions for future research that the community might consider. We hope that this position paper will provide valuable insights to the community at this pivotal stage in the advancement toward AGI. Summary of works done on RFT for MLLMs is available at https://github.com/Sun-Haoyuan23/Awesome-RL-based-Reasoning-MLLMs.

多模态强化学习推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。