将强化学习与最大似然训练统一,实现更精准的模型优化。
RL2ML: Finite-Rollout Surrogate Objectives from Reinforcement Learning to Maximum Likelihood

- 提出有限采样下的无偏梯度估计方法,保持目标与估计对齐。
- 发现更新尺度存在亚临界到超临界转变,影响训练效果。
- 最优目标选择取决于评估指标、敏感性与方差,可降维优化。
基于可验证奖励的正确性强化学习(RLVR)通过采样输出的二值反馈训练语言模型,但其期望优化目标与有限采样组带来的随机更新几何常被混淆。本文提出RL2ML,一类具有闭式精确无偏梯度估计的有限采样代理目标。该族连续连接标准强化学习、最大似然类训练及超越最大似然的目标,同时在固定采样预算下保持估计器-目标对齐。引入组级更新尺度以刻画采样组在观测到成功次数后的重加权方式,揭示了隐藏于总体目标表示中的亚临界-超临界更新尺度转变。基于此区分,校准的指标增益分析与精确方差分解表明,最佳代理目标的选择既不依赖于与最大似然的接近程度,也不仅由总体权重决定,而是联合取决于评估指标、局部敏感性和估计器方差。因此,代理目标族中剩余自由度可表述为一维优化问题,而非无约束超参数。
原文摘要 · Abstract (English)
Correctness-based Reinforcement Learning with Verifiable Rewards (RLVR) trains language models from binary feedback on sampled outputs, but the objective optimized in expectation and the stochastic update geometry induced by finite rollout groups are often conflated. This paper develops RL2ML, a family of finite-rollout surrogate objectives with a closed-form, exactly unbiased gradient estimator. The family continuously connects standard reinforcement learning, maximum-likelihood-like training, and beyond-maximum-likelihood objectives while preserving estimator-objective alignment under a fixed rollout budget. We introduce the group-level update scale to characterize how a rollout group is reweighted after its empirical success count is observed, revealing a subcritical-supercritical update-scale transition that is hidden by population-level objective notation alone. Building on this distinction, calibrated metric-gain analysis and exact variance decomposition show that the best choice of surrogate objective is determined neither by proximity to maximum likelihood nor by the population-level weight alone. Instead, it depends jointly on the evaluation metric, local sensitivity, and estimator variance. The remaining degree of freedom in the surrogate objective family can therefore be formulated as a one-dimensional optimization problem rather than treated as an unconstrained hyperparameter.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。