arXiv:2502.04270cs.LGstat.ML2025-02ICML被引 18

提出新采样策略PILAF,让模型更精准学习人类偏好。

PILAF: Optimal Human Preference Sampling for Reward Modeling

  • 通过策略插值优化反馈采样,对齐人类真实价值
  • 理论证明在优化与统计上均最优,性能超越传统方法
  • 适合需要高质量反馈的迭代与在线强化学习场景

随着大语言模型广泛应用于现实场景,使其与人类价值观对齐变得至关重要。强化学习从人类反馈(RLHF)成为关键技术,将偏好数据转化为奖励模型,以弥补无法直接获取人类真实价值的缺陷。实践中,大多数方法依赖近似奖励模型,可能无法持续引导策略向最大化原始人类价值的方向演进。本文提出政策插值对齐反馈学习(PILAF),一种新型响应采样策略,显式对齐偏好学习与最大化底层原始奖励的目标。PILAF具有理论基础,从优化与统计双重角度证明其最优性。该方法实现简单,在需要持续反馈筛选的迭代与在线RLHF设置中表现优异。

原文摘要 · Abstract (English)

As large language models increasingly drive real-world applications, aligning them with human values becomes paramount. Reinforcement Learning from Human Feedback (RLHF) has emerged as a key technique, translating preference data into reward models when oracle human values remain inaccessible. In practice, RLHF mostly relies on approximate reward models, which may not consistently guide the policy toward maximizing the underlying human values. We propose Policy-Interpolated Learning for Aligned Feedback (PILAF), a novel response sampling strategy for preference labeling that explicitly aligns preference learning with maximizing the underlying oracle reward. PILAF is theoretically grounded, demonstrating optimality from both an optimization and a statistical perspective. The method is straightforward to implement and demonstrates strong performance in iterative and online RLHF settings where feedback curation is critical.

强化学习人类对齐奖励建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。