提出新方法解决偏好优化中超参调优难题
$ξ$-DPO: Direct Preference Optimization via Ratio Reward Margin

- 用比率奖励间距重构目标,避免显式参考模型
- 新指标ξ可直接从初始奖励分布确定,无需反复试错
- 适用于需要高效调参的文本生成与对齐任务
无参考偏好优化已成为强化学习从人类反馈中获取替代方案的有效方式,其中简单偏好优化(SimPO)通过消除显式参考模型,在不依赖人工标注的情况下表现出色。然而,其核心超参数β和γ的联合调节仍是一大挑战。我们分析发现,这一困难源于SimPO中边际形式在不同奖励差距结构的数据集上缺乏可解释性。进一步研究揭示,β隐含控制样本过滤,而γ的效果依赖于数据集的奖励差距结构。基于此,我们提出ξ-DPO:通过比率奖励边际实现直接偏好优化。我们通过等价变换重构偏好目标,将优化目标从最大化奖励差距似然改为最小化奖励差距与最优边际的距离;并引入选择与拒绝响应间的比率形式奖励,有效消除了β的影响,获得有界且可解释的边际。该边际称为比率奖励边际,记为ξ。与SimPO中的γ不同,ξ明确表示选择与拒绝响应间期望的相对分离程度,可由初始奖励差距分布直接确定,从而避免重复试错调参。
原文摘要 · Abstract (English)
Reference-free preference optimization has emerged as an efficient alternative to reinforcement learning from human feedback, with Simple Preference Optimization(SimPO) demonstrating strong performance by eliminating the explicit reference model through a simple objective. However, the joint tuning of the hyperparameters $β$ and $γ$ in SimPO remains a central challenge. We argue that this difficulty arises because the margin formulation in SimPO is not easily interpretable across datasets with different reward gap structures. To better understand this issue, we conduct a comprehensive analysis of SimPO and find that $β$ implicitly controls sample filtering, while the effect of $γ$ depends on the reward gap structure of the dataset. Motivated by these observations, we propose $ξ$-DPO: Direct preference optimization via ratio reward margin. We first reformulate the preference objective through an equivalent transformation, changing the optimization target from maximizing the likelihood of reward gaps to minimizing the distance between reward gaps and optimal margins. Then, we redefine the reward in a ratio form between the chosen and rejected, which effectively cancels the effect of $β$ and yields a bounded and interpretable margin. This margin is called the ratio reward margin and is denoted by $ξ$. Unlike the margin $γ$ in SimPO, $ξ$ explicitly represents the desired relative separation between chosen and rejected responses and can be determined from the initial reward gap distribution, avoiding repeated trial-and-error tuning. ....
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。