arXiv:2602.05946cs.LGstat.ML2026-02

用散度估计统一优化大模型对齐,支持奖励与偏好双模式。

f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment

  • 基于散度估计设计新强化学习目标,统一处理奖励与偏好信号。
  • 在数学推理任务上,f-GRPO优于传统GRPO;f-HAL缓解奖励黑客问题。
  • 适合需要安全对齐但缺乏可验证奖励的场景,尤其依赖学习奖励模型时。

近期研究发现,偏好对齐目标可被解释为对齐(偏好)与非对齐(不偏好)分布间散度的估计,为设计对齐损失提供了理论基础。然而,这一视角此前仅限于基于偏好的监督。本文将其扩展至通用大模型对齐,包括可验证奖励的强化学习(RLVR),其中对齐反馈仅以标量奖励形式给出。提出$f$-组相对策略优化($f$-GRPO),一类在线策略强化学习目标,以及$f$-混合对齐损失($f$-HAL),结合在线奖励优化与离线偏好监督。证明这些目标可估计由高于/低于平均奖励响应所诱导的奖励对齐与非对齐分布之间的$f$-散度,并证明对齐后期望奖励提升。实验表明,$f$-GRPO在数学推理类RLVR任务中优于GRPO;而混合型$f$-HAL在缺乏可验证奖励、需使用学习的奖励模型时,有效缓解了奖励黑客问题。

原文摘要 · Abstract (English)

Recent work shows that preference alignment objectives can be interpreted as divergence estimators between aligned (preferred) & unaligned (less-preferred) distributions, yielding a principled recipe for designing alignment losses. However, this view has so far been limited to preference-based supervision. We extend it to general LLM alignment, including reinforcement learning with verifiable rewards (RLVR), where alignment feedback is given only as scalar rewards. We introduce $f$-Group Relative Policy Optimization ($f$-GRPO), a class of on-policy RL objectives, and $f$-Hybrid Alignment Loss ($f$-HAL), which combines on-policy reward optimization with off-policy preference supervision. We show that these objectives estimate $f$-divergences between reward-aligned & reward-unaligned distributions induced by above- & below-average reward responses, and prove expected reward improvement after alignment. Empirically, $f$-GRPO improves over GRPO on math-reasoning RLVR tasks, while hybrid $f$-HAL mitigates reward hacking in on-policy safety alignment when verifiable rewards are unavailable and learned reward models must be used.

大模型对齐强化学习奖励建模散度优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。