arXiv:2512.18215cs.LGcs.AI2025-12被引 11

提出稳定高效的单轮次强化学习框架,提升多模态模型推理能力

Stable and Efficient Single-Rollout RL for Multimodal Reasoning

  • 采用基于熵的自适应优势调节机制,防止训练崩溃
  • 仅需一半训练步数即达基线性能,相同步数下表现更优
  • 适合追求高效稳定训练的多模态推理研究者

基于可验证奖励的强化学习(RLVR)已成为提升多模态大语言模型(MLLMs)推理能力的关键范式。然而,主流的组级算法(如GRPO)需对每个提示进行多轮采样,效率低下。尽管文本领域已有单轮次变体,但在多模态场景中严重不稳定,常导致训练崩溃。为此,我们提出无组别结构的单轮次RLVR框架MSSR(Multimodal Stabilized Single-Rollout),通过熵驱动的优势调节机制自适应正则化优势幅度,有效防止崩溃并维持训练稳定。该机制在组级方法中虽已应用,但在多模态单轮次设置中并非可选,而是必要条件。在分布内评估中,MSSR实现更高训练计算效率:仅用一半训练步数即达到基线验证准确率;相同训练步数下,其性能超越基线,并在五个多样化推理密集型基准上均表现出一致的泛化提升。结果表明,MSSR实现了复杂多模态推理任务中稳定、高效且有效的强化学习优化。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) has become a key paradigm to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, prevalent group-based algorithms such as GRPO require multi-rollout sampling for each prompt. While more efficient single-rollout variants have recently been explored in text-only settings, we find that they suffer from severe instability in multimodal contexts, often leading to training collapse. To address this training efficiency-stability trade-off, we introduce $\textbf{MSSR}$ (Multimodal Stabilized Single-Rollout), a group-free RLVR framework that achieves both stable optimization and effective multimodal reasoning performance. MSSR achieves this via an entropy-based advantage-shaping mechanism that adaptively regularizes advantage magnitudes, preventing collapse and maintaining training stability. While such mechanisms have been used in group-based RLVR, we show that in the multimodal single-rollout setting they are not merely beneficial but essential for stability. In in-distribution evaluations, MSSR demonstrates superior training compute efficiency, achieving similar validation accuracy to the group-based baseline with half the training steps. When trained for the same number of steps, MSSR's performance surpasses the group-based baseline and shows consistent generalization improvements across five diverse reasoning-intensive benchmarks. Together, these results demonstrate that MSSR enables stable, compute-efficient, and effective RLVR for complex multimodal reasoning tasks.

强化学习多模态推理增强高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。