用闭式模型揭示GRPO训练动态,解释奖励上升与震荡机制。
Predictable GRPO: A Closed-Form Model of Training Dynamics

- 基于均场假设,将GRPO简化为受随机力驱动的阻尼振子。
- 预测结果与可独立测量的量相关,拟合奖励曲线决定系数达0.91以上。
- 能区分奖励劫持、优势退化等失败模式,适合算法调试者使用。
我们构建了一个基于第一性原理的降阶模型来描述训练动态。在仅依赖政策期望奖励的均场假设下,将GRPO更新简化为一个由优化器超参数和单一测得曲率尺度决定的随机驱动力阻尼振子:动量提供惯性,离策略延迟削弱阻尼,组大小主导地表现为噪声温度。该模型有三重意义:首先,其过阻尼极限涵盖经验性的单指数饱和规律,将拟合平台、时间尺度和组大小指数重新诠释为势能的固定点、逆刚度和曲率缩放指数,并通过保留惯性项捕捉到单指数无法表达的慢启动阶段;其次,预测基于可独立测量的量而非拟合参数:确定性轨迹具有组大小不变性,稳态波动为1/G,刷新间隔存在明确稳定性阈值,且呈现过阻尼到振荡的相变;第三,提供可区分多种失败模式的诊断工具——奖励劫持、优势退化、策略集中与动态不稳定性。在三个模型和两个组大小下,闭式轨迹对训练奖励的拟合决定系数达到R²≥0.91,平均轨迹在组大小上至一阶不变——不仅在奖励曲线,也在八项数学基准的分布外迁移任务中表现一致;而组内奖励差异仍残留与组大小相关的残差,超出一阶温度图景的解释范围。
原文摘要 · Abstract (English)
We develop a first-principles reduced-order model of these dynamics. Under a single mean-field assumption that summarizes the policy by its expected reward, we reduce the GRPO update to a stochastically-forced damped oscillator whose mass, damping, and stiffness are fixed in closed form by the optimizer hyperparameters together with a single measured curvature scale -- momentum supplies the inertia, off-policy lag erodes the damping, and the group size enters, to leading order, as a noise temperature. The reduction has three consequences. First, it subsumes the empirical single-exponential saturation law as its overdamped limit, recasting the fitted plateau, timescale, and size exponent as the fixed point, inverse stiffness, and curvature-scaling exponent of the underlying potential, and adding, through the retained inertial term, the slow-start phase the single exponential cannot represent. Second, it yields predictions tied to independently measurable quantities rather than fitted ones: group-size invariance of the deterministic trajectory with a $1/G$ stationary fluctuation, a sharp stability threshold in the refresh interval, and an overdamped-to-oscillatory transition. Third, it furnishes diagnostics that separate failure modes a reward curve alone conflates -- reward hacking, advantage degeneracy, policy concentration, and dynamical instability. Across three models and two group sizes, the closed-form trajectory fits training reward to $R^2 \geq 0.91$ and the mean trajectory is group-size invariant to leading order -- on both the reward curve and out-of-distribution transfer to eight math benchmarks -- while the within-group reward spread retains a residual $G$-dependence that the leading-order temperature picture does not capture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。