提出统一的Pair-GRPO框架,解决大模型对齐训练中的不稳定与梯度噪声问题。
A Unified Pair-GRPO Family: From Implicit to Explicit Preference Constraints for Stable and General RL Alignment
- 基于GRPO改进,用成对偏好替代标量奖励,保持稳定更新机制。
- 在多个基准上优于现有方法,提升对齐质量与训练稳定性,胜率超基线3.2%以上。
- 适合关注强化学习对齐、训练稳定性与可解释性的研究者和工程师。
大语言模型通过人类偏好强化学习(RLHF)对齐面临策略更新不稳定、梯度方向模糊、可解释性差及梯度方差高等问题。为此,我们建立统一的偏好驱动强化学习理论框架——Pair-GRPO家族,包含两个紧密耦合的变体:Soft-Pair-GRPO与Hard-Pair-GRPO。Soft-Pair-GRPO是对组相对策略优化(GRPO)的最小改进,将组归一化标量奖励替换为二元成对偏好奖励,保留了剪裁代理目标与KL正则结构。我们证明关键梯度等价定理:在当前策略的一阶泰勒展开下,Soft-Pair-GRPO的梯度是标准GRPO梯度的正数倍,解释了其经验稳定性。在此基础上,提出Hard-Pair-GRPO,引入显式局部概率约束与受限KL拟合优化,进一步抑制梯度噪声与全局策略漂移。为两者提供全面理论保障:单调策略提升、确定梯度方向、梯度方差降低与动态步长收敛。在标准大模型对齐基准(HH-RLHF、UltraFeedback)及连续控制任务HalfCheetah-v4上的实验表明,Pair-GRPO家族在对齐质量、人类偏好胜率、训练稳定性与泛化能力上持续优于最先进基线。消融实验验证各核心组件的关键贡献。
原文摘要 · Abstract (English)
Large language model (LLM) alignment via reinforcement learning from human preferences (RLHF) suffers from unstable policy updates, ambiguous gradient directions, poor interpretability, and high gradient variance in mainstream pairwise preference learning paradigms. To systematically address these limitations, we establish a unified theoretical framework for preference-based RL optimization centered on the Pair-GRPO family, comprising two tightly coupled variants: Soft-Pair-GRPO and Hard-Pair-GRPO. Soft-Pair-GRPO is a minimal modification of Group Relative Policy Optimization (GRPO) that replaces group-normalized scalar rewards with binary pairwise preference rewards, retaining GRPO's clipped surrogate and KL-regularized structure. We prove a critical gradient equivalence theorem: under first-order Taylor expansion around the current policy, Soft-Pair-GRPO's gradient is a positive scalar multiple of standard GRPO's gradient, explaining its empirical stability despite discarding continuous reward magnitudes. Building on this foundation, we propose Hard-Pair-GRPO, an advanced variant introducing explicit local probability constraints and constrained KL-fitting optimization to further suppress gradient noise and global policy drift. We provide comprehensive theoretical guarantees for both variants--including monotonic policy improvement, deterministic gradient direction, gradient-variance reduction, and dynamic step-size convergence. Extensive experiments on standard LLM alignment benchmarks (HH-RLHF,UltraFeedback) and the MuJoCo continuous control task HalfCheetah-v4 demonstrate that our Pair-GRPO family consistently outperforms state-of-the-art baselines in alignment quality, human preference win rate, training stability, and generalization to general reinforcement learning. Ablation studies validate the critical contributions of each core component.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。