统一视觉语言模型的监督与强化微调,提升性能并避免遗忘。
ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
- 将标签直接注入强化学习回合,实现内外双重监督
- 在多个基准上超越单独SFT、RLVR及两阶段流程
- 适合追求高效高精度视觉语言模型的开发者
后训练大型视觉-语言模型(LVLMs)通常采用监督微调(SFT)注入知识或使用可验证奖励的强化学习(RLVR)提升性能。然而,SFT常导致次优表现,而RLVR受限于模型内部知识库。虽可采用顺序SFT→RLVR流程,但计算开销大且易引发灾难性遗忘。为此,我们提出ViSurf(视觉监督与强化微调),一种统一的单阶段范式,融合SFT与RLVR优势。通过分析其训练目标,建立统一框架,将真实标签直接注入RLVR回滚过程,实现外部监督与内部强化同步进行。此外,引入三种新型奖励控制策略,保障训练稳定与优化。大量实验表明,ViSurf在多种基准上持续优于独立SFT、RLVR及传统两阶段流程。深入分析验证了其推导与设计原则的有效性。
原文摘要 · Abstract (English)
Post-training Large Vision-and-Language Models (LVLMs) typically involves Supervised Fine-Tuning (SFT) for knowledge injection or Reinforcement Learning with Verifiable Rewards (RLVR) for performance enhancement. However, SFT often leads to sub-optimal performance, while RLVR remains constrained by the model's internal knowledge base. While a sequential SFT $\rightarrow$ RLVR pipeline can be used, it introduces significant computational overhead and suffers from catastrophic forgetting. To address these limitations, we propose ViSurf (\textbf{Vi}sual \textbf{Su}pervised-and-\textbf{R}einforcement \textbf{F}ine-Tuning), a unified, single-stage paradigm that integrates the strengths of both SFT and RLVR. By analyzing their training objectives, we establish a unified framework that injects ground-truth labels directly into RLVR rollouts, facilitating simultaneous external supervision and internal reinforcement. Furthermore, we introduce three novel reward control strategies to ensure training stability and optimization. Extensive experiments demonstrate that ViSurf consistently outperforms standalone SFT, RLVR, and the traditional two-stage pipeline across diverse benchmarks. In-depth analysis corroborates these findings, validating the derivation and design principles of ViSurf.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。