arXiv:2606.09059cs.LGcs.AI2026-06

Stage-1主要影响模型熵分布,而非最终性能。

Stage-1 Controls the Entropy Regime, Not the Outcome

  • 通过对比SFT与OPD初始化,发现熵水平差异显著
  • 仅在初始阶段提升答案多样性,对最终效果无显著影响
  • 适合关注训练过程动态与熵控制的研究者

两阶段后训练(先阶段1暖启动,再阶段2强化学习)在视觉语言模型中日益流行。本研究以Qwen2.5-VL-7B为实验模型,使用同模态72B VLM教师进行基于策略的蒸馏(OPD),在小数据条件下分析阶段1的作用。三种暖启动方法在Geometry3K内部验证上均落在53%–54%的狭窄区间,表明阶段1对域内终点性能影响有限。而采用匹配配方的早停SFT使域外MathVista得分提升+2.1点,逆转了过训练版本的-9.5点下降。最明显区别在于熵状态:OPD进入强化学习时政策熵显著高于SFT初始化,且该差异贯穿整个训练轨迹。在域内初始化阶段,OPD还表现出更高的答案多样性与pass@16(较SFT高出+2.0至+5.2点),但问题级置信区间显示差异不显著。强化学习后,终点pass@16值差距不超过1.1点,且在MathVista上六种模型差异均在1.2点以内。结论是:阶段1主要关联于熵状态,其下游收益微小、局部化,不足以证明OPD优于SFT作为强化学习起点。

原文摘要 · Abstract (English)

Two-stage post-training -- a Stage-1 warm-start (supervised fine-tuning, SFT, or on-policy distillation, OPD) followed by Stage-2 reinforcement learning (RL) -- is increasingly used for vision-language models (VLMs). We ask what Stage-1 actually controls in a small-data study using Qwen2.5-VL-7B with a same-modality 72B VLM teacher for OPD. First, the three warm-starts reach a narrow $53$--$54\%$ band on Geometry3K internal validation, consistent with the narrow range reported by recent specialized methods; this setup provides little evidence that Stage-1 changes the in-domain endpoint. Second, a matched-recipe, early-stopped SFT improves out-of-domain MathVista by $+2.1$ points, reversing the $-9.5$-point drop of an over-trained variant. The clearest difference is the \emph{entropy regime}: OPD enters RL with substantially higher policy entropy than either SFT initialization, and the separation remains visible through the available trajectories. At the in-domain initialization, OPD also has higher answer diversity and pass@16 ($+2.0$ to $+5.2$ points over SFT), although problem-level bootstrap intervals show that the smaller contrast is uncertain. The advantage is absent after RL (endpoint pass@16 values within $1.1$ points) and on MathVista (six models within $1.2$ points). Our contribution is therefore a bounded empirical characterization: Stage-1 is strongly associated with the entropy regime in this setup, but the downstream payoff is small, localized, and not evidence that OPD is a better RL warm-start.

强化学习视觉语言模型熵控制训练阶段

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。