arXiv:2512.12690cs.LGcs.CL2025-12被引 3

重新评估监督微调在视觉语言模型推理中的作用,发现其在小模型和数据效率上表现更优。

Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning

  • 通过控制变量对比监督微调与强化学习,发现效果依赖模型能力与数据规模。
  • 2000条数据的监督微调性能可媲美甚至超过2万条数据的强化学习。
  • 监督微调在跨模态迁移中表现更强,且避免强化学习中奖励欺骗问题。

近期视觉语言模型(VLM)推理能力的提升主要归因于强化学习(RL),导致学界对监督微调(SFT)的关注减弱。许多研究认为引入SFT不仅无助于提升推理能力,反而可能损害训练效果。本文通过系统、受控的实验对比SFT与RL在VLM推理中的表现。在使用相同数据源的前提下,我们发现两者相对有效性取决于模型容量、数据规模与数据分布。与普遍假设相反,研究结果表明:(1)在较弱模型中,SFT更可靠地激发推理能力;(2)仅用2000条数据的SFT即可达到或优于使用2万条数据的RL表现;(3)SFT在跨模态迁移中具备更强泛化性。此外,我们发现强化学习存在普遍的奖励欺骗问题——更高奖励并不对应更高推理准确率。这些结果挑战了“强化学习优于监督微调”的主流观点,强调应重新审视SFT的价值,并倡导在后训练阶段将SFT与RL作为互补组件协同使用。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) reasoning have been largely attributed to the rise of reinforcement Learning (RL), which has shifted the community's focus away from the supervised fine-tuning (SFT) paradigm. Many studies suggest that introducing the SFT stage not only fails to improve reasoning ability but may also negatively impact model training. In this study, we revisit this RL-centric belief through a systematic and controlled comparison of SFT and RL on VLM Reasoning. Using identical data sources, we find that the relative effectiveness of SFT and RL is conditional and strongly influenced by model capacity, data scale, and data distribution. Contrary to common assumptions, our findings show that SFT plays a crucial role across several scenarios: (1) Effectiveness for weaker models. SFT more reliably elicits reasoning capabilities in smaller or weaker VLMs. (2) Data efficiency. SFT with only 2K achieves comparable or better reasoning performance to RL with 20K. (3) Cross-modal transferability. SFT demonstrates stronger generalization across modalities. Moreover, we identify a pervasive issue of deceptive rewards, where higher rewards fail to correlate with better reasoning accuracy in RL. These results challenge the prevailing "RL over SFT" narrative. They highlight that the role of SFT may have been underestimated and support a more balanced post-training pipeline in which SFT and RL function as complementary components.

视觉语言模型监督微调强化学习推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。