arXiv:2507.22607cs.CVcs.AI2025-07被引 26

用渐进式强化学习提升多模态推理能力

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning

  • 分阶段逐步提升任务难度,引导模型高效学习
  • 动态调整难度与推理路径长度,平衡效率与准确率
  • 在数学、科学等多领域表现优于现有模型

强化学习已被证明能有效提升大语言模型的推理能力。近期研究将这一范式逐步扩展至多模态推理任务。由于多模态任务在语义内容和问题形式上固有的复杂性和多样性,现有模型在不同领域和难度水平下常表现出性能不稳定。为解决这一问题,我们提出VL-Cogito,一个通过新型多阶段渐进式课程强化学习(PCuRL)框架训练的先进多模态推理模型。PCuRL系统性地引导模型经历难度逐级提升的任务,显著提升其在多样化多模态场景中的推理能力。该框架引入两项关键创新:(1) 在线难度软加权机制,动态调节连续强化学习阶段的训练难度;(2) 动态长度奖励机制,促使模型根据任务复杂度自适应调整推理路径长度,从而在推理效率与正确性之间取得平衡。实验评估表明,VL-Cogito在涵盖数学、科学、逻辑和通用理解的主流多模态基准上持续匹配或超越现有推理导向模型,验证了方法的有效性。

原文摘要 · Abstract (English)

Reinforcement learning has proven its effectiveness in enhancing the reasoning capabilities of large language models. Recent research efforts have progressively extended this paradigm to multimodal reasoning tasks. Due to the inherent complexity and diversity of multimodal tasks, especially in semantic content and problem formulations, existing models often exhibit unstable performance across various domains and difficulty levels. To address these limitations, we propose VL-Cogito, an advanced multimodal reasoning model trained via a novel multi-stage Progressive Curriculum Reinforcement Learning (PCuRL) framework. PCuRL systematically guides the model through tasks of gradually increasing difficulty, substantially improving its reasoning abilities across diverse multimodal contexts. The framework introduces two key innovations: (1) an online difficulty soft weighting mechanism, dynamically adjusting training difficulty across successive RL training stages; and (2) a dynamic length reward mechanism, which encourages the model to adaptively regulate its reasoning path length according to task complexity, thus balancing reasoning efficiency with correctness. Experimental evaluations demonstrate that VL-Cogito consistently matches or surpasses existing reasoning-oriented models across mainstream multimodal benchmarks spanning mathematics, science, logic, and general understanding, validating the effectiveness of our approach.

多模态推理强化学习渐进学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。