arXiv:2508.17000cs.CLcs.LG2025-08被引 3

提出一种新型强化学习方法,让大模型生成更符合人类反馈。

KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF

  • 从动作价值角度设计基于KL正则的在线强化学习算法
  • 在摘要和对话任务上性能媲美PPO,胜率更高
  • 为语言模型对齐提供新思路,适合研究人机对齐者

近端策略优化(PPO)是当前语言模型强化学习从人类反馈(LM-RLHF)中广泛使用的一种有效策略梯度算法。尽管其表现良好,但其动机具有启发性,且对LM-RLHF中使用的KL散度约束处理方式较为随意。本文提出一种新的动作价值强化学习方法——KL正则化Q学习(KLQ),用于LM-RLHF场景。我们证明该方法在特定意义上等价于某种形式的PPO,尽管其出发点完全不同。我们在两个关键语言生成任务——摘要生成与单轮对话——上对KLQ进行了基准测试。结果表明,KLQ在优化LM-RLHF目标上与PPO表现相当,并在大模型作为裁判的评估中始终取得更高的胜率。

原文摘要 · Abstract (English)

Proximal Policy Optimisation (PPO) is an established and effective policy gradient algorithm used for Language Model Reinforcement Learning from Human Feedback (LM-RLHF). PPO performs well empirically but has a heuristic motivation and handles the KL-divergence constraint used in LM-RLHF in an ad-hoc manner. In this paper, we develop a a new action-value RL method for the LM-RLHF setting, KL-regularised Q-Learning (KLQ). We then show that our method is equivalent to a version of PPO in a certain specific sense, despite its very different motivation. Finally, we benchmark KLQ on two key language generation tasks -- summarisation and single-turn dialogue. We demonstrate that KLQ performs on-par with PPO at optimising the LM-RLHF objective, and achieves a consistently higher win-rate against PPO on LLM-as-a-judge evaluations.

强化学习大模型对齐语言生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。