arXiv:2507.14843cs.LGcs.AI2025-07被引 68

RLVR提升精准度但可能困住模型,限制新解法探索

The Invisible Leash: Why RLVR May or May Not Escape Its Origin

  • 将RLVR视为受基础模型分布约束的优化机制,限制新解发现
  • 大采样预算下,正确答案的可访问性反而下降,支持集收缩更严重
  • 看似更不确定的生成路径,最终收敛到更少的固定答案

近期进展表明,基于可验证奖励的强化学习(RLVR)是提升大语言模型能力的有前景方法。然而,当前实践是否真正拓展了模型的推理边界,还是仅放大了基础模型已知的高奖励输出以提高精度,尚不明确。本研究通过实证分析揭示了RLVR的局限性:其作为支持约束优化机制,可能受限于基础模型初始分布,难以发现全新解法。我们发现熵与奖励存在权衡——尽管RLVR稳定提升 exttt{pass@1},但会逐步缩小探索范围,可能忽略正确但低频的解。大量实验表明,在更大采样预算下,经验支持集的收缩普遍超过扩张,导致原本可访问的正确答案无法恢复。有趣的是,虽然局部熵上升,但生成步级不确定性增加,答案级熵下降,说明看似更随机的路径最终收敛至更少的确定答案。这揭示了RLVR在扩展推理边界上的潜在瓶颈。突破这一‘隐形缰绳’需未来创新,主动向低频解空间注入概率质量。

原文摘要 · Abstract (English)

Recent advances highlight Reinforcement Learning with Verifiable Rewards (RLVR) as a promising method for enhancing LLMs' capabilities. However, it remains unclear whether the current practice of RLVR truly expands a model's reasoning boundary or mainly amplifies high-reward outputs that the base model already knows, thereby improving precision. This study presents an empirical investigation that provides fresh insights into the limits of RLVR. We examine how RLVR can operate as a support-constrained optimization mechanism that may restrict the discovery of entirely original solutions, remaining constrained by the base model's initial distribution. We also identify an entropy-reward trade-off: while RLVR reliably enhances precision, it may progressively narrow exploration and potentially overlook correct yet underrepresented solutions. Extensive empirical experiments validate that while RLVR consistently improves \texttt{pass@1}, \textit{the shrinkage of empirical support generally outweighs the expansion of empirical support under larger sampling budgets}, failing to recover correct answers that were previously accessible to the base model. Interestingly, while RLVR sometimes increases token-level entropy, it results in greater uncertainty at each generation step and declining answer-level entropy. This indicates that these seemingly more uncertain paths ultimately converge onto a smaller set of distinct answers. Taken together, we reveal potential limits of RLVR in extending reasoning horizons. Breaking this invisible leash requires future innovations that seed probability mass into underrepresented solution regions.

强化学习推理能力语言模型可验证奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。