arXiv:2410.15610cs.LG2024-10被引 2

首次证明神经网络参数化下在线RLHF的全局收敛性,为大模型对齐提供理论保障。

On The Global Convergence Of Online RLHF With Neural Parametrization

  • 采用双层优化框架与弱梯度支配假设,设计一阶求解算法。
  • 实现样本复杂度ε^{-7/2}的收敛速率,达到当前最优。
  • 理论成果适用于真实大模型场景,适合关注对齐安全的研究者。

强化学习从人类反馈(RLHF)在对齐大语言模型与人类价值观方面至关重要。该过程包含监督微调、奖励学习和策略学习三个阶段。尽管已有离线与在线对齐方法,但普遍存在分布偏移问题,源于无法准确捕捉奖励学习与策略学习阶段间的分布依赖关系。现有方法多为近似处理,理论分析仅限于表格型设置,难以推广至实际场景。本文基于Kwon等(2024)提出的双层框架,结合弱梯度支配假设,首次在神经网络参数化设置下建立RLHF的全局收敛性,获得ε^{-7/2}的样本复杂度。核心贡献包括:(i) 提出参数化设置下的双层优化框架,并提出一阶求解方法;(ii) 分析算法理论收敛率,推导出当前最优界。据我们所知,这是首个在神经网络参数化设置下建立收敛率与全局最优性的RLHF工作。

原文摘要 · Abstract (English)

The importance of Reinforcement Learning from Human Feedback (RLHF) in aligning large language models (LLMs) with human values cannot be overstated. RLHF is a three-stage process that includes supervised fine-tuning (SFT), reward learning, and policy learning. Although there are several offline and online approaches to aligning LLMs, they often suffer from distribution shift issues. These issues arise from the inability to accurately capture the distributional interdependence between the reward learning and policy learning stages. Consequently, this has led to various approximated approaches, but the theoretical insights and motivations remain largely limited to tabular settings, which do not hold in practice. This gap between theoretical insights and practical implementations is critical. It is challenging to address this gap as it requires analyzing the performance of AI alignment algorithms in neural network-parameterized settings. Although bi-level formulations have shown promise in addressing distribution shift issues, they suffer from the hyper-gradient problem, and current approaches lack efficient algorithms to solve this. In this work, we tackle these challenges employing the bi-level formulation laid out in Kwon et al. (2024) along with the assumption \emph{Weak Gradient Domination} to demonstrate convergence in an RLHF setup, obtaining a sample complexity of $ε^{-\frac{7}{2}}$ . Our key contributions are twofold: (i) We propose a bi-level formulation for AI alignment in parameterized settings and introduce a first-order approach to solve this problem. (ii) We analyze the theoretical convergence rates of the proposed algorithm and derive state-of-the-art bounds. To the best of our knowledge, this is the first work to establish convergence rate bounds and global optimality for the RLHF framework in neural network-parameterized settings.

RLHF强化学习大模型对齐收敛性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。