arXiv:2606.19744cs.CLcs.AI2026-06

研究顺序优化对人类偏好学习的影响,发现遗忘模式不统一。

Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settings

论文配图:Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settings
图 1 · 摘自论文原文
  • 分阶段用DPO优化不同偏好目标,考察其相互影响
  • 不同目标组合下,偏好变化从部分退化到正向迁移不等
  • 适合关注模型对齐中多目标冲突与兼容性的研究者

将语言模型对齐人类偏好通常需要优化多个行为目标。实际中常采用直接偏好优化(DPO)按顺序应用这些目标,但尚不清楚后期训练是否均匀损害早期学习的偏好,或其影响是否依赖于目标间关系。我们针对四种偏好设置——分布冲突、多属性交互、强安全信号和兼容响应质量——研究了顺序DPO的影响。使用Llama-3.1-8B-Instruct搭配LoRA适配器,每阶段后以固定基线模型为参考评估所有目标。结果表明,顺序DPO不会产生单一遗忘模式:偏好变化范围从部分退化、稳定、成对重分配到正向迁移,取决于目标关系、信号强度与训练顺序。通过长度归一化的策略边际进行成对分析显示,聚合指标可能掩盖偏好对间的异质性变化;四分位分解揭示高置信度对在不同设置下可能退化或提升。机制诊断表明,在所有设置中,第二阶段梯度与适配器更新与前一目标近乎正交,缺乏直接梯度对抗的有力证据。这些发现提示,未来顺序对齐流程应考虑目标兼容性与信号强度,而非默认后期目标会均匀影响前期偏好。

原文摘要 · Abstract (English)

Aligning language models with human preferences often requires optimising multiple behavioural objectives. A practical approach is to apply these objectives sequentially using preference optimisation methods such as Direct Preference Optimisation (DPO), but it remains unclear whether later training uniformly degrades preferences learned earlier or whether the effect depends on the relationship between objectives. We study sequential DPO across four preference settings covering distributional conflict, multi-attribute interaction, strong safety signal, and compatible response-quality objectives. Using Llama-3.1-8B-Instruct with LoRA adapters, we evaluate all objectives after every stage with a fixed base-model reference. We find that sequential DPO does not produce a single forgetting pattern; preference change ranges from partial degradation to stability, pair-level redistribution, or positive transfer depending on objective relationship, signal strength, and training order. Pair-level analysis using length-normalised policy margins shows that aggregate metrics can mask heterogeneous changes across preference pairs, whereas quartile decomposition reveals that high-confidence pairs can either degrade or improve depending on the setting. Mechanistic diagnostics show that Stage~2 gradients and adapter updates are near-orthogonal to the previous objective across all settings, providing little evidence that direct gradient opposition is the primary driver. These findings suggest that future sequential alignment pipelines should account for objective compatibility and signal strength, rather than assuming that later objectives affect earlier preferences uniformly.

偏好优化模型对齐多目标学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。