用小模型做探索,让大模型学得更好更快。
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO

- 用固定小模型生成有结构的探索轨迹,避免随机噪声。
- 在AIME24上提升8.8%准确率,且减少计算开销。
- 适合需要高效训练大模型的推理任务研究者。
我们发现,在相同模型家族中,较小模型天然具备更高的策略级多样性,表现为随样本量增加时,其pass@k性能优于更大模型。与引入逐标记随机性不同,这种多样性具有时间相关性,保持逻辑连贯性,可为梯度估计提供结构化探索信号。为此提出S2L-PO框架:以固定小模型作为自然探索者,指导大模型训练。通过渐进退火策略,从离线小模型采样平滑过渡到大模型自主采样,避免因小模型能力限制导致的训练中性能下降,实现更快收敛和更高性能上限。S2L-PO在多个数学推理基准上表现优异(如使用1.7B小模型引导8B模型时,AIME24准确率提升8.8%),同时降低推演计算成本。
原文摘要 · Abstract (English)
We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollouts, prevailing strategies primarily increase diversity by injecting more token-level randomness, which may introduce step-wise noise and lead to incoherent trajectories. We uncover that smaller models within the same model family inherently exhibit higher policy-level diversity, indicated by their superior pass@k relative to larger counterparts as sample counts increase. Unlike token-level noise, this diversity is temporally correlated, preserves logical consistency, and provides structured exploration signals for gradient estimation. We thus propose S2L-PO (Small-to-Large Policy Optimization), a framework that leverages fixed small models as natural explorers to train larger models. To balance exploration and exploitation, we design a progressive annealing strategy that transitions from offline small-model rollouts to the large learner's own sampling. This shift elegantly avoids mid-training performance drops caused by the small model's capacity limits, achieving faster convergence and unlocking a higher performance ceiling. S2L-PO improves accuracy on diverse mathematical reasoning benchmarks (e.g., +8.8% on AIME 24 using a 1.7B explorer to guide the 8B model) while reducing rollout compute.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。