用不完美的奖励模型加速在线强化学习,提升效率。
Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective
- 基于策略覆盖性理论,利用偏差但相关的奖励模型迁移知识
- 提出TPO算法,理论证明比标准方法更高效
- 可适配DPO等方法,实测在摘要任务中提升性能
样本效率对在线人类反馈强化学习(RLHF)至关重要。尽管已有研究关注高效探索策略,但利用存在偏差但相关性强的奖励模型以加速学习的潜力仍被低估。本文从RLHF目标中KL正则化的特性出发,发现策略对最优策略的覆盖能力与其次优性直接相关。基于此,提出新的迁移学习原则和理论算法——转移策略优化(TPO),在理论上优于标准在线学习。受理论启发,设计基于胜率的策略选择机制,提升计算效率。该经验性迁移方法具有模块化特性,可与DPO、IPO、XPO等多种策略优化方法结合,进一步提升性能。在摘要生成任务上的实验验证了方法的有效性。
原文摘要 · Abstract (English)
Sample efficiency is critical for online Reinforcement Learning from Human Feedback (RLHF). While existing works investigate sample-efficient online exploration strategies, the potential of utilizing misspecified yet relevant reward models to accelerate learning remains underexplored. This paper studies how to transfer knowledge from those imperfect reward models in online RLHF. We start by identifying a novel property due to KL-regularization in the RLHF objective: \emph{a policy's coverability of the optimal policy is captured by its sub-optimality}. Building on this insight, we propose novel transfer learning principles and a theoretical algorithm -- \emph{\textbf{T}ransfer \textbf{P}olicy \textbf{O}ptimization (\textbf{TPO})} -- with provable benefits compared to standard online learning. Empirically, inspired by our theoretical findings, we develop a win-rate-based transfer policy selection strategy with improved computational efficiency. Moreover, our empirical transfer learning technique is modular and can be integrated with various policy optimization methods, such as DPO, IPO and XPO, to further enhance their performance. We validate the effectiveness of our method through experiments on summarization tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。