用Pass@k训练提升大模型推理中的探索能力
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

- 以Pass@k为奖励信号训练模型,增强探索性
- 相比Pass@1,能有效避免陷入局部最优
- 适合研究大模型强化学习与推理优化的学者
基于可验证奖励的强化学习(RLVR)通常采用Pass@1作为奖励,但面临探索与利用难以平衡的问题,导致策略偏好保守行为,收敛至局部最优。尽管先前工作在评估中使用过Pass@k,但其与大语言模型探索能力在RLVR中的关联仍被忽视。本文首次将Pass@k作为奖励用于策略训练(即Pass@k Training),观察到模型探索能力显著提升。进一步推导出该方法的优势解析解,实现高效训练。分析表明,探索与利用并非固有冲突,反而可相互促进。此外,Pass@k Training本质上是直接设计优势函数,由此启发我们初步探索RLVR中的优势函数设计,取得良好效果,揭示了未来重要方向。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards (RLVR), which typically adopts Pass@1 as the reward, has faced the issues in balancing exploration and exploitation, causing policies to prefer conservative actions, converging to a local optimum. Identifying an appropriate reward metric is therefore crucial. Regarding the prior work, although Pass@k has been used in evaluation, its connection to LLM exploration ability in RLVR remains largely overlooked. To investigate this, we first use Pass@k as the reward to train the policy model (i.e., $\textbf{Pass@k Training}$), and observe the improvement on its exploration ability. Next, we derive an analytical solution for the advantage of Pass@k Training, leading to an efficient and effective process. Building on this, our analysis reveals that exploration and exploitation are not inherently conflicting objectives, while they can mutually enhance each other. Moreover, Pass@k Training with analytical derivation essentially involves directly designing the advantage function. Inspired by this, we preliminarily explore the advantage design for RLVR, showing promising results and highlighting a potential future direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。