arXiv:2410.02504stat.MLcs.LG2024-10被引 21

通过智能选对话和老师,用最少人力提升大模型对齐效果

Dual Active Learning for Reinforcement Learning from Human Feedback

  • 同时选择最值得评价的对话和最合适的人类教师
  • 在有限标注预算下,奖励估计误差最小化,策略偏差随样本量平方根下降
  • 适合需要高效利用人工反馈的AI对齐研究者

对齐大语言模型(LLMs)与人类偏好是生成式人工智能近期进展的关键。强化学习从人类反馈(RLHF)被广泛用于实现此目标,其核心在于从人类反馈中学习奖励函数。然而,人类反馈成本高、耗时长,因此需高效收集高质量对话数据供人工标注。此外,不同标注者专业水平不一,如何选择最优教师至关重要。本文将对齐问题建模为离线强化学习任务,受$D$-最优设计启发,提出一种双通道主动奖励学习算法,实现对话与教师的协同优化选择。进一步采用悲观强化学习求解对齐问题,基于学习到的奖励估计器。理论上,我们证明所提自适应选择策略可使奖励估计器的广义方差渐近最小,并证明悲观策略的次优性以$O(1/\sqrt{T})$随给定样本预算$T$衰减。仿真及在大模型上的实验表明,该算法显著优于现有方法。

原文摘要 · Abstract (English)

Aligning large language models (LLMs) with human preferences is critical to recent advances in generative artificial intelligence. Reinforcement learning from human feedback (RLHF) is widely applied to achieve this objective. A key step in RLHF is to learn the reward function from human feedback. However, human feedback is costly and time-consuming, making it essential to collect high-quality conversation data for human teachers to label. Additionally, different human teachers have different levels of expertise. It is thus critical to query the most appropriate teacher for their opinions. In this paper, we use offline reinforcement learning (RL) to formulate the alignment problem. Motivated by the idea of $D$-optimal design, we first propose a dual active reward learning algorithm for the simultaneous selection of conversations and teachers. Next, we apply pessimistic RL to solve the alignment problem, based on the learned reward estimator. Theoretically, we show that the reward estimator obtained through our proposed adaptive selection strategy achieves minimal generalized variance asymptotically, and prove that the sub-optimality of our pessimistic policy scales as $O(1/\sqrt{T})$ with a given sample budget $T$. Through simulations and experiments on LLMs, we demonstrate the effectiveness of our algorithm and its superiority over state-of-the-arts.

RLHF主动学习大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。