提出新型损失函数,让模型对齐更稳定且理论更可靠。
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
- 发现非凸f散度也能保持对齐可解性,放宽了传统限制。
- 设计新损失避免概率坍塌,提升训练稳定性。
- 新方法理论更强、实测表现不输经典DPO,适合追求稳健性的研究者。
DPO及其相关算法通过直接优化强化学习人类反馈(RLHF)目标来对齐语言模型:在最大化Bradley-Terry奖励的同时,通过KL散度惩罚使策略贴近参考策略。先前工作表明,即使将KL散度替换为具有凸生成函数f的f-散度家族,原问题仍可解。本文第一项贡献是证明f的凸性并非必要,提出了更一般的“DPO诱导”条件,精确刻画了何时该问题仍可解。第二项贡献是建立一个关于f的额外条件,以防止概率位移这一已知的实证现象——即优劣响应的概率同时趋近于零。满足此条件的f称为抗位移型(displacement-resistant)。最后,我们聚焦于一个同时满足DPO诱导与抗位移特性的具体f,提出新型平方形式的SquaredPO损失。相比DPO,该损失提供更强的理论保障,且在实践中表现相当。
原文摘要 · Abstract (English)
DPO and related algorithms align language models by directly optimizing the RLHF objective: find a policy that maximizes the Bradley-Terry reward while staying close to a reference policy through a KL divergence penalty. Previous work showed that this approach could be further generalized: the original problem remains tractable even if the KL divergence is replaced by a family of $f$-divergence with a convex generating function $f$. Our first contribution is to show that convexity of $f$ is not essential. Instead, we identify a more general condition, referred to as DPO-inducing, that precisely characterizes when the RLHF problem remains tractable. Our next contribution is to establish a second condition on $f$ that is necessary to prevent probability displacement, a known empirical phenomenon in which the probabilities of the winner and the loser responses approach zero. We refer to any $f$ that satisfies this condition as displacement-resistant. We finally focus on a specific DPO-inducing and displacement-resistant $f$, leading to our novel SquaredPO loss. Compared to DPO, this new loss offers stronger theoretical guarantees while performing competitively in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。