arXiv:2503.07065cs.CV2025-03被引 79

小模型也能强推理:用分阶段强化学习提升视觉语言模型泛化能力

Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning

  • 分阶段强化学习,由易到难逐步训练模型能力
  • 30亿参数模型性能媲美320亿参数大模型
  • 适合资源有限但需强推理的小模型研究者

当前先进视觉语言模型(VLMs)虽在复杂多模态任务中表现卓越,但其性能高度依赖大规模模型扩展,限制了实际部署。小规模VLMs虽更实用,但在传统监督微调下面临域外泛化和推理能力不足的问题,远逊于当代大语言模型。为此,我们提出课程强化微调(Curr-ReFT),一种专为小规模VLMs设计的后训练范式。该方法包含两个阶段:(1) 课程强化学习,通过难度感知奖励机制,从基础视觉感知逐步过渡到复杂推理任务;(2) 基于拒绝采样的自提升,通过精选高质量多模态与语言样本保留基础能力。大量实验表明,采用Curr-ReFT训练的模型在多种视觉任务中均达到领先性能,且30亿参数模型表现堪比320亿参数模型,证明高效训练范式可有效缩小小模型与大模型间的差距。

原文摘要 · Abstract (English)

While state-of-the-art vision-language models (VLMs) have demonstrated remarkable capabilities in complex visual-text tasks, their success heavily relies on massive model scaling, limiting their practical deployment. Small-scale VLMs offer a more practical alternative but face significant challenges when trained with traditional supervised fine-tuning (SFT), particularly in two aspects: out-of-domain (OOD) generalization and reasoning abilities, which significantly lags behind the contemporary Large language models (LLMs). To address these challenges, we propose Curriculum Reinforcement Finetuning (Curr-ReFT), a novel post-training paradigm specifically designed for small-scale VLMs. Inspired by the success of reinforcement learning in LLMs, Curr-ReFT comprises two sequential stages: (1) Curriculum Reinforcement Learning, which ensures steady progression of model capabilities through difficulty-aware reward design, transitioning from basic visual perception to complex reasoning tasks; and (2) Rejected Sampling-based Self-improvement, which maintains the fundamental capabilities of VLMs through selective learning from high-quality multimodal and language examples. Extensive experiments demonstrate that models trained with Curr-ReFT paradigm achieve state-of-the-art performance across various visual tasks in both in-domain and out-of-domain settings. Moreover, our Curr-ReFT enhanced 3B model matches the performance of 32B-parameter models, demonstrating that efficient training paradigms can effectively bridge the gap between small and large models.

视觉语言模型强化学习小模型推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。