统一轨迹聚合方式,让强化学习自动选择稳定或激进的更新策略
One Ring to Rule Them All: Unifying Group-Based RL via Dynamic Power-Mean Geometry
- 用幂均几何指数p统一平均方式,动态调节更新强度
- 在数学推理任务上超越现有强基线,提升训练稳定性与性能
- 适合需要自适应策略优化的复杂强化学习场景
基于群体的强化学习从算术平均(GRPO)发展到几何平均(GMPO),但二者均依赖固定聚合方式,忽略轨迹演化的异质性。本文提出幂均策略优化(PMPO),通过参数化幂均几何指数p,统一前述方法。理论上,调整p可调控梯度更新集中度,实现基于优势贡献的权重重分配。为自适应确定p,提出剪裁感知有效样本量(ESS)机制:根据轨迹剪裁比例设定目标ESS,求解对应p值。使PMPO能动态切换——对可靠轨迹采用激进算术平均,对不稳定轨迹采用保守几何平均。在多个数学推理基准测试中,PMPO显著优于强基线。
原文摘要 · Abstract (English)
Group-based reinforcement learning has evolved from the arithmetic mean of GRPO to the geometric mean of GMPO. While GMPO improves stability by constraining a conservative objective, it shares a fundamental limitation with GRPO: reliance on a fixed aggregation geometry that ignores the evolving and heterogeneous nature of each trajectory. In this work, we unify these approaches under Power-Mean Policy Optimization (PMPO), a generalized framework that parameterizes the aggregation geometry via the power-mean geometry exponent p. Within this framework, GRPO and GMPO are recovered as special cases. Theoretically, we demonstrate that adjusting p modulates the concentration of gradient updates, effectively reweighting tokens based on their advantage contribution. To determine p adaptively, we introduce a Clip-aware Effective Sample Size (ESS) mechanism. Specifically, we propose a deterministic rule that maps a trajectory clipping fraction to a target ESS. Then, we solve for the specific p to align the trajectory induced ESS with this target one. This allows PMPO to dynamically transition between the aggressive arithmetic mean for reliable trajectories and the conservative geometric mean for unstable ones. Experiments on multiple mathematical reasoning benchmarks demonstrate that PMPO outperforms strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。