arXiv:2503.06810cs.LGcs.AI2025-03被引 7

用悲观策略防止强化学习中偏好过拟合,提升模型对人类偏好的真实响应能力。

Mitigating Preference Hacking in Policy Optimization with Pessimism

  • 引入悲观目标函数,利用不确定性来抑制过优化行为。
  • 在摘要和助手任务中表现稳定,显著降低偏好欺骗现象。
  • 适合关注对齐安全、防止模型‘钻空子’的研究者使用。

本文针对基于人类反馈的强化学习(RLHF)中的过优化问题提出解决方案。该技术依赖于固定偏好数据集训练的奖励或偏好模型,但这些模型在数据分布外评估时不可靠,导致常见的奖励或偏好欺骗现象。本文提出新的悲观化目标函数,通过不确定性下的悲观策略,理论上保证对过优化的鲁棒性,并设计了实用算法P3O与PRPO以优化这些目标。该方法适用于一般偏好优化场景,也可用于奖励模型。在文档摘要和构建帮助型助手的任务中,验证了P3O与PRPO对过优化具有显著抗性。

原文摘要 · Abstract (English)

This work tackles the problem of overoptimization in reinforcement learning from human feedback (RLHF), a prevalent technique for aligning models with human preferences. RLHF relies on reward or preference models trained on \emph{fixed preference datasets}, and these models are unreliable when evaluated outside the support of this preference data, leading to the common reward or preference hacking phenomenon. We propose novel, pessimistic objectives for RLHF which are provably robust to overoptimization through the use of pessimism in the face of uncertainty, and design practical algorithms, P3O and PRPO, to optimize these objectives. Our approach is derived for the general preference optimization setting, but can be used with reward models as well. We evaluate P3O and PRPO on the tasks of fine-tuning language models for document summarization and creating helpful assistants, demonstrating remarkable resilience to overoptimization.

强化学习偏好对齐鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。