arXiv:2503.16081cs.LGcs.IR2025-03被引 23

用动态强化学习提升多模态大模型的泛化推理能力

OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning

  • 设计动态KL策略的分组相对策略优化方法,突破传统强化学习瓶颈
  • 在同任务评估中比SFT提升超5.72%,跨任务评估平均提升超61.63%
  • 适合追求多模态模型泛化能力的研究者与开发者

多模态大语言模型(MLLM)在处理多种输入数据并生成上下文相关输出方面表现出色。尽管监督微调(SFT)是主流任务优化方法,但难以培养泛化推理能力。强化学习(RL)虽具潜力,却面临两大挑战:其一,多模态任务中的泛化能力尚未充分探索;其二,训练中常使用固定KL散度或钳制策略,导致性能瓶颈。为此,我们提出OThink-MR1,一种具备深度理解与推理能力的先进MLLM。核心是引入动态KL策略的分组相对策略优化(GRPO-D),显著提升强化学习性能。在Qwen2-VL-2B-Instruct上,GRPO-D在两个适配数据集上的同任务评估中,相比SFT提升超过5.72%,相比GRPO提升超过13.59%。此外,在跨任务评估中,其平均相对提升超过61.63%。结果表明,基于GRPO-D训练的模型可在不同多模态任务间有效迁移,凸显了OThink-MR1卓越的泛化推理能力。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have gained significant traction for their ability to process diverse input data types and generate coherent, contextually relevant outputs across various applications. While supervised fine-tuning (SFT) has been the predominant approach to enhance MLLM capabilities in task-specific optimization, it often falls short in fostering crucial generalized reasoning abilities. Although reinforcement learning (RL) holds great promise in overcoming these limitations, it encounters two significant challenges: (1) its generalized capacities in multimodal tasks remain largely unexplored, and (2) its training constraints, including the constant Kullback-Leibler divergence or the clamp strategy, often result in suboptimal bottlenecks. To address these challenges, we propose OThink-MR1, an advanced MLLM equipped with profound comprehension and reasoning capabilities across multimodal tasks. Specifically, we introduce Group Relative Policy Optimization with a dynamic Kullback-Leibler strategy (GRPO-D), which markedly enhances reinforcement learning (RL) performance. For Qwen2-VL-2B-Instruct, GRPO-D achieves a relative improvement of more than 5.72% over SFT and more than 13.59% over GRPO in same-task evaluation on two adapted datasets. Furthermore, GRPO-D demonstrates remarkable cross-task generalization capabilities, with an average relative improvement of more than 61.63% over SFT in cross-task evaluation. These results highlight that the MLLM trained with GRPO-D on one multimodal task can be effectively transferred to another task, underscoring the superior generalized reasoning capabilities of our proposed OThink-MR1 model.

多模态强化学习泛化推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。