提出统一理论框架,解释为何强化学习对齐更优。
Beyond RLHF: A Unified Theoretical Framework of Alignment
- 将对齐建模为基于成对偏好分布学习
- 三类新目标均实现1/n收敛率,避免退化
- 首次解释策略类方法为何优于似然方法
基于人类反馈的强化学习(RLHF)已成为控制大语言模型输出质量的主流范式。然而现有理论无法充分证明RLHF目标的有效性,且不同方法常在不同框架下分析,难以比较保障性。为此,本文将对齐重新定义为从成对偏好中学习分布,引入概率假设描述偏好如何揭示目标语言模型信息。由此提出三种原则性对齐目标:偏好最大似然估计、偏好提炼和反KL最小化。证明三者均具有强非渐近收敛性,收敛速率为O(1/n),自然避免退化。其中反KL与RLHF高度相似,为后者提供坚实理论依据。此外,该理论首次解释了经验上发现的“在线策略目标(如RLHF)通常优于似然型目标(如DPO)”的现象。实验表明,所提目标在多个任务与模型上均达到强基线水平。
原文摘要 · Abstract (English)
Alignment via reinforcement learning from human feedback (RLHF) has become the dominant paradigm for controlling the quality of outputs from large language models (LLMs). However, existing theories do not provide strong justification for the RLHF objective itself and do not allow comparisons of the guarantees between various methods because different methods are often analyzed under different frameworks. Toward a unified framework for alignment, we ask under what assumptions can we derive existing or new training objectives and obtain theoretical guarantees. To this end, we reframe alignment as distribution learning from pairwise preferences, which makes a probabilistic assumption describing how preferences reveal information about the target LM. This leads us to propose three principled alignment objectives: preference maximum likelihood estimation, preference distillation, and reverse KL minimization. We prove that they all enjoy strong non-asymptotic $O(1/n)$ convergence to the target LM, naturally avoiding degeneracy. In particular, reverse KL highly resembles the RLHF objective, providing strong justification for RLHF. Furthermore, our theory explains, for the first time, the empirical finding that on-policy objectives (e.g., RLHF) typically outperform likelihood-style objectives (e.g., DPO). Finally, empirical results indicate that the proposed objectives are competitive with strong baselines across several tasks and models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。