提出SPO框架,解决偏好学习中的灾难性偏好偏移问题
Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
- 基于双层优化构建稳定偏好优化框架,约束对齐区域防止偏好偏移
- 在SimPO上将推理准确率从37.5%提升至接近原始水平,避免性能崩溃
- 适用于需要可靠对齐的高风险场景,如医疗、金融领域模型训练
直接偏好学习已成为主流的离线偏好优化范式。多数方法基于布拉德利-特瑞(BT)模型进行成对偏好排序,直接对齐语言模型与人类偏好。先前研究观察到一种反直觉现象——似然位移,即在训练过程中优选响应的绝对概率同时下降。我们证明这种位移可能导致更严重的失败模式,称为‘灾难性偏好偏移’,即丢失的偏好概率质量意外转移到分布外(OOD)响应上。这一失败模式是所有基于BT的方法共有的关键限制,源于判别对齐与生成基础能力之间的根本冲突,最终导致严重性能退化(例如,SimPO的推理准确率从73.5%降至37.5%)。我们从概率演化视角分析现有BT方法,理论证明其过度依赖模型初始化并引发偏好偏移。为此,我们提出一个理论驱动的稳定偏好优化(SPO)框架,将偏好学习限制在安全对齐区域内。实证评估表明,SPO有效稳定并提升了现有BT类方法的性能。SPO为偏好学习目标设计提供了新洞见,开辟了更可靠、可解释的语言模型对齐新路径。
原文摘要 · Abstract (English)
Direct Preference Learning has emerged as a dominant offline paradigm for preference optimization. Most of these methods are based on the Bradley-Terry (BT) model for pairwise preference ranking, which directly aligns language model with human preference. Prior work has observed a counter-intuitive phenomenon termed likelihood displacement, where the absolute probability of preferred responses decreases simultaneously during training. We demonstrate that such displacement can lead to a more devastating failure mode, which we defined as \textit{Catastrophic Preference Shift}, where the lost preference probability mass inadvertently shifts toward out-of-distribution (OOD) responses. Such a failure mode is a key limitation shared across BT-style direct preference learning methods, due to the fundamental conflict between the unconstrained discriminative alignment and generative foundational capabilities, ultimately leading to severe performance degradation (e.g., SimPO suffers a significant drop in reasoning accuracy from 73.5\% to 37.5\%). We analyze existing BT-style methods from the probability evolution perspective and theoretically prove that these methods exhibit over-reliance on model initialization and can lead to preference shift. To resolve these counter-intuitive behaviors, we propose a theoretically grounded Stable Preference Optimization (SPO) framework that constrains preference learning within a safe alignment region. Empirical evaluations demonstrate that SPO effectively stabilizes and enhances the performance of existing BT-style preference learning methods. SPO provides new insights into the design of preference learning objectives and opens up new avenues towards more reliable and interpretable language model alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。