让大模型对不完美的奖励模型更抗干扰,提升对齐稳定性。
Reward-Robust RLHF in LLMs
- 用贝叶斯奖励模型集成建模奖励不确定性,平衡性能与鲁棒性。
- 在多个基准上优于基线,准确率更高且长期学习更稳定。
- 适合关注模型对齐可靠性与抗奖励欺骗的研究者和开发者。
随着大语言模型向更高级智能演进,基于人类反馈的强化学习(RLHF)被视为实现通用人工智能(AGI)的关键路径。然而,依赖奖励模型(RM)进行对齐的方法存在固有不稳定性和缺陷,易引发奖励黑客和与人类意图的错位。本文提出一种奖励鲁棒的RLHF框架,通过引入贝叶斯奖励模型集成(BRME),建模奖励函数的不确定性集,优化目标兼顾名义性能与最小奖励信号,从而在不完美奖励模型下仍能实现稳定学习。实验表明,该框架在多个基准上持续优于基线,提升准确率并增强长期稳定性。理论分析进一步证明,该方法可逼近恒定奖励设定下的稳定性,即使在随机情形下也表现可接受。这些成果展示了其在提升大模型对齐性能与稳定性方面的潜力。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) continue to progress toward more advanced forms of intelligence, Reinforcement Learning from Human Feedback (RLHF) is increasingly seen as a key pathway toward achieving Artificial General Intelligence (AGI). However, the reliance on reward-model-based (RM-based) alignment methods introduces significant challenges due to the inherent instability and imperfections of Reward Models (RMs), which can lead to critical issues such as reward hacking and misalignment with human intentions. In this paper, we introduce a reward-robust RLHF framework aimed at addressing these fundamental challenges, paving the way for more reliable and resilient learning in LLMs. Our approach introduces a novel optimization objective that carefully balances performance and robustness by incorporating Bayesian Reward Model Ensembles (BRME) to model the uncertainty set of reward functions. This allows the framework to integrate both nominal performance and minimum reward signals, ensuring more stable learning even with imperfect RMs. Empirical results demonstrate that our framework consistently outperforms baselines across diverse benchmarks, showing improved accuracy and long-term stability. We also provide a theoretical analysis, demonstrating that reward-robust RLHF approaches the stability of constant reward settings, which proves to be acceptable even in a stochastic-case analysis. Together, these contributions highlight the framework potential to enhance both the performance and stability of LLM alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。