用行为策略信息修正偏好学习,让强化学习更贴近人类反馈
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
- 用后悔值建模人类偏好,显式捕捉行为策略信息
- 在高维连续控制任务中,离线与在线设置均显著提升性能
- 适合关注奖励对齐与策略优化的强化学习研究者
为设计符合人类目标的奖励函数,基于人类反馈的强化学习(RLHF)已成为从人类偏好中学习奖励函数并利用强化学习算法优化策略的重要方法。然而,现有RLHF方法常将轨迹误认为由最优策略生成,导致似然估计不准和学习效果欠佳。受直接偏好优化框架启发,该文提出策略标注的偏好学习(PPL),通过引入反映行为策略信息的后悔值来建模人类偏好,解决似然不匹配问题。同时,基于后悔值原理推导出对比式KL正则化,以增强序列决策中的RLHF表现。在高维连续控制任务中的实验表明,PPL在离线和在线设置下均显著提升性能。
原文摘要 · Abstract (English)
To design rewards that align with human goals, Reinforcement Learning from Human Feedback (RLHF) has emerged as a prominent technique for learning reward functions from human preferences and optimizing policies via reinforcement learning algorithms. However, existing RLHF methods often misinterpret trajectories as being generated by an optimal policy, causing inaccurate likelihood estimation and suboptimal learning. Inspired by Direct Preference Optimization framework which directly learns optimal policy without explicit reward, we propose policy-labeled preference learning (PPL), to resolve likelihood mismatch issues by modeling human preferences with regret, which reflects behavior policy information. We also provide a contrastive KL regularization, derived from regret-based principles, to enhance RLHF in sequential decision making. Experiments in high-dimensional continuous control tasks demonstrate PPL's significant improvements in offline RLHF performance and its effectiveness in online settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。