用分布奖励模型统一解决强化学习中的奖励不确定性问题
A Unifying Lens on Reward Uncertainty in RLHF
- 提出分布式奖励模型,通过KL正则化构建统一的悲观奖励机制
- 推导出闭式有效奖励公式,理论涵盖多种现有启发式方法
- 为奖励模型集成提供统一框架,适合研究者与工业部署者参考
基于人类反馈的强化学习(RLHF)受限于奖励滥用问题,即策略利用代理奖励模型(RM)的误差,获得高分但无真实质量提升。一种自然缓解方式是悲观性:在奖励模型不确定区域降低奖励。然而,标准标量奖励模型无法提供不确定性的合理度量。本文认为正确对象应为分布奖励模型 $p(r arget x,y)$。在贝叶斯推断或KL分布鲁棒优化(KL-DRO)视角下,KL正则化的RLHF目标可导出闭式有效奖励 $ ilde r(x,y) = \eta\log\mathbb{E}_p[e^{\pm r/\beta}]$。悲观分支统一了已有奖励模型集成启发式:均值聚合、最坏情况优化(WCO)和不确定性加权优化(UWO)皆为此表达式的极限或截断形式。该结果也揭示了各方法的隐含假设。
原文摘要 · Abstract (English)
Reinforcement learning from human feedback (RLHF) is bottlenecked by reward hacking, where the policy exploits errors in a proxy reward model (RM) and produces high RM scores without genuine quality gains. A natural mitigation is pessimism: lowering rewards in regions where the RM is uncertain. However, standard scalar RMs provide no principled notion of uncertainty. We argue that the right object is a distributional reward model $p(r\mid x,y)$. Under either a Bayesian inference or a KL-distributionally robust optimization (KL-DRO) lens, the KL-regularized RLHF objective admits a closed-form effective reward $\tilde r(x,y) = \pmβ\log\mathbb{E}_p[e^{\pm r/β}]$. The pessimistic branch unifies the prior heuristics for RM ensemble aggregation: mean aggregation, worst-case optimization (WCO), and uncertainty-weighted optimization (UWO) all emerge as limits or truncations of this single expression. This also clarifies the implicit assumptions of each existing rule.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。