用散度估计统一优化大模型对齐,支持奖励与偏好双模式。
f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment
- 基于散度估计设计新强化学习目标,统一处理奖励与偏好信号。
- 在数学推理任务上,f-GRPO优于传统GRPO;f-HAL缓解奖励黑客问题。
- 适合需要安全对齐但缺乏可验证奖励的场景,尤其依赖学习奖励模型时。
近期研究发现,偏好对齐目标可被解释为对齐(偏好)与非对齐(不偏好)分布间散度的估计,为设计对齐损失提供了理论基础。然而,这一视角此前仅限于基于偏好的监督。本文将其扩展至通用大模型对齐,包括可验证奖励的强化学习(RLVR),其中对齐反馈仅以标量奖励形式给出。提出$f$-组相对策略优化($f$-GRPO),一类在线策略强化学习目标,以及$f$-混合对齐损失($f$-HAL),结合在线奖励优化与离线偏好监督。证明这些目标可估计由高于/低于平均奖励响应所诱导的奖励对齐与非对齐分布之间的$f$-散度,并证明对齐后期望奖励提升。实验表明,$f$-GRPO在数学推理类RLVR任务中优于GRPO;而混合型$f$-HAL在缺乏可验证奖励、需使用学习的奖励模型时,有效缓解了奖励黑客问题。
原文摘要 · Abstract (English)
Recent work shows that preference alignment objectives can be interpreted as divergence estimators between aligned (preferred) & unaligned (less-preferred) distributions, yielding a principled recipe for designing alignment losses. However, this view has so far been limited to preference-based supervision. We extend it to general LLM alignment, including reinforcement learning with verifiable rewards (RLVR), where alignment feedback is given only as scalar rewards. We introduce $f$-Group Relative Policy Optimization ($f$-GRPO), a class of on-policy RL objectives, and $f$-Hybrid Alignment Loss ($f$-HAL), which combines on-policy reward optimization with off-policy preference supervision. We show that these objectives estimate $f$-divergences between reward-aligned & reward-unaligned distributions induced by above- & below-average reward responses, and prove expected reward improvement after alignment. Empirically, $f$-GRPO improves over GRPO on math-reasoning RLVR tasks, while hybrid $f$-HAL mitigates reward hacking in on-policy safety alignment when verifiable rewards are unavailable and learned reward models must be used.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。