提出最优奖励建模方法,减少人类偏好标注成本。
Optimal Design for Reward Modeling in RLHF
- 用线性上下文双臂老虎机框架优化偏好数据选择。
- 在嵌入空间线性假设下,理论证明了简单遗憾上界。
- 首个提供离线训练与最坏情况保证的理论工作。
基于人类反馈的强化学习(RLHF)已成为对齐语言模型与人类偏好的主流方法。该方法需收集大量文本生成的成对人类偏好数据,并据此推断(显式或隐式)奖励模型。尽管已有多种方法用于学习奖励模型并实现对齐,但高昂的人类偏好标注成本尚未得到充分关注,亟需理论指导。本文针对此问题,形式化了RLHF中的奖励建模过程,将有效数据集的选择建模为简单的后悔最小化任务,采用线性上下文双臂老虎机方法。考虑到可能存在的大量候选选项,该方法比最优臂识别更具一致性。随后提出一个离线求解框架,在奖励模型在嵌入空间中线性且奖励参数有界的假设下,导出了简单遗憾的理论边界。最后,给出了与上界仅差常数和对数项的下界。据我们所知,这是该领域首个提供离线方法及最坏情况保证的理论贡献。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) has become a popular approach to align language models (LMs) with human preferences. This method involves collecting a large dataset of human pairwise preferences across various text generations and using it to infer (implicitly or explicitly) a reward model. Numerous methods have been proposed to learn the reward model and align a LM with it. However, the costly process of collecting human preferences has received little attention and could benefit from theoretical insights. This paper addresses this issue and aims to formalize the reward training model in RLHF. We frame the selection of an effective dataset as a simple regret minimization task, using a linear contextual dueling bandit method. Given the potentially large number of arms, this approach is more coherent than the best-arm identification setting. We then propose an offline framework for solving this problem. Under appropriate assumptions - linearity of the reward model in the embedding space, and boundedness of the reward parameter - we derive bounds on the simple regret. Finally, we provide a lower bound that matches our upper bound up to constant and logarithmic terms. To our knowledge, this is the first theoretical contribution in this area to provide an offline approach as well as worst-case guarantees.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。