arXiv:2505.19731stat.MLcs.LG2025-05被引 5

用博弈论方法更稳定地从人类反馈中训练语言模型。

Proximal Point Nash Learning from Human Feedback

  • 将人类偏好建模为博弈,通过近端点框架优化
  • 证明算法在高概率下收敛,且比传统方法更稳定
  • 适合需要可靠反馈学习的大型语言模型微调

传统基于人类反馈的强化学习(RLHF)常依赖奖励模型,假设偏好符合布拉德利-特瑞模型,但该假设难以刻画真实人类偏好的复杂性(如非传递性)。纳什学习从人类反馈(NLHF)提供了一种更直接的替代方案,将问题视为寻找由偏好定义的博弈的纳什均衡。尽管已有研究直接在策略空间中求解纳什学习问题,本文考虑更具现实意义的策略参数化设置。我们首先分析一种简单的自对弈策略梯度方法,其等价于在线 IPO。我们建立了该方法的高概率逐次迭代收敛性保证,但分析也揭示了其底层动态可能存在的稳定性局限。受此启发,我们将自对弈更新嵌入近端点框架,得到一个稳定的算法。对于该组合方法,我们证明了高概率逐次迭代收敛性,并讨论其更实用的版本——纳什近端(Nash Prox)。最后,我们将该方法应用于大语言模型的后训练,并验证了其经验性能。

原文摘要 · Abstract (English)

Traditional Reinforcement Learning from Human Feedback (RLHF) often relies on reward models, frequently assuming preference structures like the Bradley--Terry model, which may not accurately capture the complexities of real human preferences (e.g., intransitivity). Nash Learning from Human Feedback (NLHF) offers a more direct alternative by framing the problem as finding a Nash equilibrium of a game defined by these preferences. While many works study the Nash learning problem directly in the policy space, we instead consider it under a more realistic policy parametrization setting. We first analyze a simple self-play policy gradient method, which is equivalent to Online IPO. We establish high-probability last-iterate convergence guarantees for this method, but our analysis also reveals a possible stability limitation of the underlying dynamics. Motivated by this, we embed the self-play updates into a proximal point framework, yielding a stabilized algorithm. For this combined method, we prove high-probability last-iterate convergence and discuss its more practical version, which we call Nash Prox. Finally, we apply this method to post-training of large language models and validate its empirical performance.

强化学习人类反馈纳什均衡语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。