RLVR通过强化可解步骤概率,让大模型从已有能力中涌现复杂推理。
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
- 用概率框架定义能力,聚焦单步操作优化
- 多步任务成功率由各步骤概率联合决定,相关性达0.69~0.96
- 适合研究模型涌现机制或强化学习优化策略的学者
强化学习结合可验证奖励(RLVR)是否赋予大语言模型新能力,还是仅激发潜在能力,仍是核心争议。本文支持前者观点,提出一种概率框架,将能力定义为实例级可解性。我们假设复杂推理的出现源于原子步骤概率的增强,从而克服多步推理中成功概率指数衰减的难题。利用Algebrarium框架,我们在仅单步操作上训练模型,并评估其在未见多步任务上的表现。实验结果表明:(1) RLVR通过放大模型现有技能,激励探索此前不可达的解题路径;(2) 复合性能严格受原子步骤联合概率支配,皮尔逊相关系数ρ∈[0.69, 0.96];(3) RLVR作为全局优化器,可能牺牲特定技能以最大化整体奖励。本工作为RLVR中的涌现能力提供了新解释,表明对可解问题的迭代优化,使模型具备解决先前不可解场景的能力。
原文摘要 · Abstract (English)
Whether Reinforcement Learning with Verifiable Rewards (RLVR) endows Large Language Models (LLMs) with new capabilities or merely elicits latent traces remains a central debate. In this work, we align with the former view, proposing a probabilistic framework where capability is defined by instance-level solvability. We hypothesize that the emergence of complex reasoning can be driven by sharpening atomic step probabilities, which enables models to overcome the exponential decay of success rates inherent in multi-step reasoning chains. Utilizing the Algebrarium framework, we train models exclusively on single-step operations and evaluate their performance on unseen multi-step tasks. Our empirical results confirm that: (1) RLVR incentivizes the exploration of previously inaccessible solution paths by amplifying the model's existing skills; (2) composite performance is strictly governed by the joint probability of atomic steps, evidenced by high Pearson correlation coefficients ($ρ\in [0.69, 0.96]$); and (3) RLVR, acting as a global optimizer, can cause specific skills to be sacrificed to maximize aggregate reward. Our work offers a novel explanation for emergent abilities in RLVR, suggesting that the iterative optimization of solvable problems enables models to develop the capabilities to tackle previously unsolvable scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。