arXiv:2510.13651cs.LGmath.OC2025-10被引 16

揭示大模型强化学习中二元奖励算法的本质是概率梯度上升。

What is the objective of reasoning with reinforcement learning?

  • 将多种强化学习算法视为单调变换后的正确答案概率梯度上升。
  • 拒绝采样对应对数变换,GRPO对应平方根的反正弦变换。
  • 为大模型推理优化提供统一理论视角,适合研究强化学习机制者阅读。

我们证明,在大型语言模型中使用二元奖励的几种流行强化学习算法,可被视作对给定提示下正确答案概率的单调变换进行随机梯度上升。特别是,与拒绝采样相关的变换是其对数,而与GRPO算法相关的变换是平方根的反正弦函数。该分析揭示了不同算法在优化目标上的统一本质,为理解基于奖励的大模型推理提供了新的理论框架。

原文摘要 · Abstract (English)

We show that several popular algorithms for reinforcement learning in large language models with binary rewards can be viewed as stochastic gradient ascent on a monotone transform of the probability of a correct answer given a prompt. In particular, the transformation associated with rejection sampling algorithms is the logarithm and that associated with the GRPO algorithm is the arcsine of the square root.

强化学习大模型推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。