提出一种新型强化学习方法,让大模型生成更符合人类反馈。
KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
- 从动作价值角度设计基于KL正则的在线强化学习算法
- 在摘要和对话任务上性能媲美PPO,胜率更高
- 为语言模型对齐提供新思路,适合研究人机对齐者
近端策略优化(PPO)是当前语言模型强化学习从人类反馈(LM-RLHF)中广泛使用的一种有效策略梯度算法。尽管其表现良好,但其动机具有启发性,且对LM-RLHF中使用的KL散度约束处理方式较为随意。本文提出一种新的动作价值强化学习方法——KL正则化Q学习(KLQ),用于LM-RLHF场景。我们证明该方法在特定意义上等价于某种形式的PPO,尽管其出发点完全不同。我们在两个关键语言生成任务——摘要生成与单轮对话——上对KLQ进行了基准测试。结果表明,KLQ在优化LM-RLHF目标上与PPO表现相当,并在大模型作为裁判的评估中始终取得更高的胜率。
原文摘要 · Abstract (English)
Proximal Policy Optimisation (PPO) is an established and effective policy gradient algorithm used for Language Model Reinforcement Learning from Human Feedback (LM-RLHF). PPO performs well empirically but has a heuristic motivation and handles the KL-divergence constraint used in LM-RLHF in an ad-hoc manner. In this paper, we develop a a new action-value RL method for the LM-RLHF setting, KL-regularised Q-Learning (KLQ). We then show that our method is equivalent to a version of PPO in a certain specific sense, despite its very different motivation. Finally, we benchmark KLQ on two key language generation tasks -- summarisation and single-turn dialogue. We demonstrate that KLQ performs on-par with PPO at optimising the LM-RLHF objective, and achieves a consistently higher win-rate against PPO on LLM-as-a-judge evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。