从贝叶斯视角揭示DPO奖励设计的内在逻辑与训练动态本质
Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
- 将偏好优化视为学习政策更新所需的差分信息分布
- 发现高熵差分信息提升指令跟随能力,低熵则利于知识问答
- 为基于偏好的对齐提供理论基础与实践指导
直接偏好优化(DPO)被广泛用于以监督方式对齐语言模型与人类偏好。然而,其对数比率奖励的合理性、偏好数据集的统计结构如何影响训练动态,以及这些动态如何影响下游能力等问题仍未解决。本文从贝叶斯视角出发,将偏好优化的目标定义为学习将参考策略更新为目标策略所需的差分信息。为此,我们提出差分信息分布(DID),即携带贝叶斯证据以更新策略的样本分布。通过分析DID,我们获得三个互补洞见:首先,当偏好编码了更新策略所需差分信息时,DPO的对数比率奖励具有唯一合理性;其次,常见于DPO的训练动态——如对数似然变化和策略探索——源于幂律形式的DID关系;最后,我们利用DID的熵作为学习信息不确定性的合理度量,发现高熵DID有助于开放性指令遵循,而低熵DID有利于知识密集型问答。结果表明,DPO的奖励设计、训练动态与下游性能均是学习差分信息的自然结果,为基于偏好的对齐提供了理论基础与实践指引。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) has been widely used for aligning language models with human preferences in a supervised manner. However, several key questions remain unresolved: the rationale behind its log-ratio reward, how the statistical structure of preference datasets shapes its training dynamics, and how those dynamics impact downstream capabilities. We approach these questions from a Bayesian perspective, interpreting the goal of preference optimization as learning the differential information required to update a reference policy into a target policy. To formalize this view, we introduce the Differential Information Distribution (DID), defined as the distribution over samples that carry the Bayesian evidence required to update policies. We introduce three complementary insights by viewing preference optimization through the DID. First, we find that DPO's log-ratio reward is uniquely justified when preferences encode the Differential Information needed to update a reference policy into the target policy. Second, we discuss how commonly observed training dynamics in DPO, including changes in log-likelihood and policy exploration, stem from a power-law DID relationship. Finally, we analyze how training dynamics influence downstream performance using the entropy of DID, a principled measure of uncertainty in the learned information. We observe that learning high-entropy DID improves open-ended instruction-following, while low-entropy DID benefits knowledge-intensive QA. Taken together, our results show that DPO's reward design, training dynamics, and downstream capabilities all emerge as natural consequences of learning Differential Information, offering both a principled theoretical foundation and practical guidance for preference-based alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。