arXiv:2501.18873cs.LG2025-01被引 2

基于轨迹偏好反馈,更稳健地找到最优策略。

Best Policy Learning from Trajectory Preference Feedback

  • 用后验采样方法直接从偏好比较中学习策略
  • 在模拟和图像生成任务上优于现有基线方法
  • 适合需要优化生成模型的交互式场景

基于人类反馈的强化学习(RLHF)虽强大,但依赖学习到的奖励模型,易受误设和奖励黑客影响。基于偏好的强化学习(PbRL)通过直接利用轨迹间的二元偏好比较,提供更鲁棒的替代方案。本文研究了PbRL中的最优策略识别问题,旨在优化生成模型,如多轮交互后的后训练阶段。该学习设定结合了离线偏好数据集(可能有偏差或分布外)与在线纯探索,因此系统性在线学习至关重要。为此,我们提出后验采样偏好学习(PSPL),一种受双顶汤普森采样启发的新算法,对奖励模型和动态模型保持后验分布。我们首次为PbRL提供了贝叶斯简单遗憾保证,并引入高效近似方法,在仿真和图像生成基准上表现优于现有基线。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful approach for aligning generative models, but its reliance on learned reward models makes it vulnerable to mis-specification and reward hacking. Preference-based Reinforcement Learning (PbRL) offers a more robust alternative by directly leveraging noisy binary comparisons over trajectories. We study the best policy identification problem in PbRL, motivated by post-training optimization of generative models, for example, during multi-turn interactions. Learning in this setting combines an offline preference dataset - potentially biased or out-of-distribution and collected from a rater of subpar `competence' - with online pure exploration, making systematic online learning essential. To this end, we propose Posterior Sampling for Preference Learning ($\mathsf{PSPL}$), a novel algorithm inspired by Top-Two Thompson Sampling that maintains posteriors over the reward model and dynamics. We provide the first Bayesian simple regret guarantees for PbRL and introduce an efficient approximation that outperforms existing baselines on simulation and image generation benchmarks.

强化学习偏好学习生成模型策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。