arXiv:2412.09544cs.LGcs.AI2024-12被引 12

提出新方法缓解偏好优化中的奖励欺骗问题,提升对齐效果。

Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking

  • 结合加权熵与鲁棒奖励最大化,对抗统计波动导致的奖励偏差。
  • 动态调整偏好标签,使不可靠样本梯度逐渐消失,提升模型稳定性。
  • 在AlpacaEval 2.0等基准上优于DPO,最高提升13.0分,适合对齐研究者使用。

AI系统对齐人类偏好常受奖励欺骗困扰,尤其在离线偏好优化中更为显著。本文识别出两类由数据集统计波动引发的奖励欺骗:Ⅰ型因劣质选择被误判为更优,Ⅱ型因优质选择被误判为较差。证明多数主流或理论性偏好优化方法均受此影响。为此,提出POWER方法,融合Guiasu加权熵与鲁棒奖励最大化,在一般函数逼近下具备有限样本保证,性能媲美数据中最佳策略。针对Ⅱ型欺骗,分析学习动态并设计动态标签更新机制,使偏好标签趋向稳定状态,降低不可靠样本梯度。实验表明,结合动态标签的POWER-DL在多个对齐基准上持续领先,相较DPO在AlpacaEval 2.0上提升最高达13.0点,Arena-Hard上提升11.5点,同时保持或提升数学推理等下游任务性能。理论与实证均验证了其缓解奖励欺骗的有效性。

原文摘要 · Abstract (English)

Aligning AI systems with human preferences typically suffers from the infamous reward hacking problem, where optimization of an imperfect reward model leads to undesired behaviors. In this paper, we investigate reward hacking in offline preference optimization, which aims to improve an initial model using a preference dataset. We identify two types of reward hacking stemming from statistical fluctuations in the dataset: Type I Reward Hacking due to subpar choices appearing more favorable, and Type II Reward Hacking due to decent choices appearing less favorable. We prove that many (mainstream or theoretical) preference optimization methods suffer from both types of reward hacking. To mitigate Type I Reward Hacking, we propose POWER, a new preference optimization method that combines Guiasu's weighted entropy with a robust reward maximization objective. POWER enjoys finite-sample guarantees under general function approximation, competing with the best covered policy in the data. To mitigate Type II Reward Hacking, we analyze the learning dynamics of preference optimization and develop a novel technique that dynamically updates preference labels toward certain "stationary labels", resulting in diminishing gradients for untrustworthy samples. Empirically, POWER with dynamic labels (POWER-DL) consistently outperforms state-of-the-art methods on alignment benchmarks, achieving improvements of up to 13.0 points on AlpacaEval 2.0 and 11.5 points on Arena-Hard over DPO, while also improving or maintaining performance on downstream tasks such as mathematical reasoning. Strong theoretical guarantees and empirical results demonstrate the promise of POWER-DL in mitigating reward hacking.

对齐优化奖励欺骗偏好学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。