arXiv:2504.02587cs.LGcs.CL2025-04被引 18

提出可复现的视觉语言模型强化学习框架,解决训练难比、结果难评问题。

Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme

  • 从零构建四步透明强化学习流程,支持多模型多数据集
  • 发现响应长度受随机种子影响,反思行为与输出长度相关
  • 强化学习在泛化上优于监督微调,即使数据质量高也更优

强化学习(RL)近年来在提升大语言模型推理能力方面展现出强大潜力,正被积极拓展至视觉语言模型(VLMs)。然而,现有VLM中的RL应用多依赖高度定制化框架,阻碍可复现性与可及性,且缺乏标准化评估协议,难以比较结果或解析训练动态。本文提出一个透明的、从头开始的VLM强化学习框架,提供一个最小但功能完整的四步流程,在多个模型和数据集上验证有效。同时,提出标准化评估方案以衡量训练动态与反思行为。在视觉推理任务上的大量实验揭示关键实证发现:响应长度对随机种子敏感,反思行为与输出长度相关,且即便使用高质量数据,强化学习在泛化性能上始终优于监督微调(SFT)。这些发现结合所提框架,旨在建立可复现基线,推动基于强化学习的VLM研究广泛参与。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has recently shown strong potential in improving the reasoning capabilities of large language models and is now being actively extended to vision-language models (VLMs). However, existing RL applications in VLMs often rely on heavily engineered frameworks that hinder reproducibility and accessibility, while lacking standardized evaluation protocols, making it difficult to compare results or interpret training dynamics. This work introduces a transparent, from-scratch framework for RL in VLMs, offering a minimal yet functional four-step pipeline validated across multiple models and datasets. In addition, a standardized evaluation scheme is proposed to assess training dynamics and reflective behaviors. Extensive experiments on visual reasoning tasks uncover key empirical findings: response length is sensitive to random seeds, reflection correlates with output length, and RL consistently outperforms supervised fine-tuning (SFT) in generalization, even with high-quality data. These findings, together with the proposed framework, aim to establish a reproducible baseline and support broader engagement in RL-based VLM research.

强化学习视觉语言模型可复现性评估标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。