统一了大模型偏好学习的理论框架,帮用户选对方法
From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models
- 从三个维度解析不同对齐方法的本质差异:偏好模型、正则化机制、数据分布
- 揭示在线与离线方法在覆盖率上的根本区别,给出过优化的缩放规律
- 提供可落地的选择指南,解决长度劫持、模式崩溃等常见失败问题
大语言模型与人类偏好的对齐已成为安全有益AI部署的关键。尽管基于人类反馈的强化学习(RLHF)是主流范式,但近年来涌现大量替代方法——直接偏好优化(DPO)、身份偏好优化(IPO)、卡尼曼-特沃斯基优化(KTO)、简单偏好优化(SimPO)等,令实践者难以抉择。本文提供偏好学习的理论统一视角,表明看似多样实则源于三个正交轴上的合理选择:(I) 偏好模型(目标函数背后的似然模型),(II) 正则化机制(如何控制与参考策略的偏差),(III) 数据分布(在线/离线学习及覆盖要求)。通过精确定义与定理推导,我们得出关键结果:在线与离线方法存在覆盖率分离,奖励过优化存在缩放规律,且直接对齐方法在特定设计组合下会失效。分析揭示长度劫持、模式崩溃、似然位移等失败现象由特定组合引发。综合50余篇实证研究,提出实用决策指南,将偏好学习从经验艺术转变为理论驱动的学科。
原文摘要 · Abstract (English)
Aligning large language models (LLMs) with human preferences has become essential for safe and beneficial AI deployment. While Reinforcement Learning from Human Feedback (RLHF) established the dominant paradigm, a proliferation of alternatives -- Direct Preference Optimization (DPO), Identity Preference Optimization (IPO), Kahneman-Tversky Optimization (KTO), Simple Preference Optimization (SimPO), and many others -- has left practitioners without clear guidance on method selection. This survey provides a \textit{theoretical unification} of preference learning methods, revealing that the apparent diversity reduces to principled choices along three orthogonal axes: \textbf{(I) Preference Model} (what likelihood model underlies the objective), \textbf{(II) Regularization Mechanism} (how deviation from reference policies is controlled), and \textbf{(III) Data Distribution} (online vs.\ offline learning and coverage requirements). We formalize each axis with precise definitions and theorems, establishing key results including the coverage separation between online and offline methods, scaling laws for reward overoptimization, and conditions under which direct alignment methods fail. Our analysis reveals that failure modes -- length hacking, mode collapse, likelihood displacement -- arise from specific, predictable combinations of design choices. We synthesize empirical findings across 50+ papers and provide a practitioner's decision guide for method selection. The framework transforms preference learning from an empirical art into a theoretically grounded discipline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。