arXiv:2506.19780cs.LG2025-06被引 2

让AI同时理解多种偏好,自动调整权重,更贴近真实人类判断。

Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing

  • 用多维偏好混合建模,突破单一目标限制。
  • 在6个基准上优于基线,小模型也表现稳健。
  • 自适应权重调度,避免人工设定偏差,适合多场景对齐。

基于直接偏好优化(DPO)的对齐方法将偏好学习重构为成对比较的监督优化,相比强化学习从人类反馈(RLHF)更具效率和稳定性。然而,现有DPO方法隐含假设单一固定偏好目标,难以捕捉现实人类判断中多维度且常冲突的特性。本文提出列表式直接偏好优化(λ-DPO),统一提升监督粒度与偏好灵活性。不同于将多维偏好压缩为单一排序,λ-DPO 构建由偏好向量λ加权的列表偏好分布混合,使单个模型内化连续的偏好权衡谱。为进一步增强鲁棒性,引入基于性能驱动的随机λ调度器,根据下游实测性能自适应采样偏好权重,显式缓解静态权重方案固有的误设风险。我们在多个模型族与规模上,于六个常用基准上评估该方法,实验结果表明其持续优于基线。

原文摘要 · Abstract (English)

Recent alignment methods based on Direct Preference Optimization (DPO) reformulate preference learning as supervised optimization over pairwise comparisons, offering improved efficiency and stability over reinforcement learning from human feedback (RLHF). However, existing DPO-style methods implicitly assume a single fixed preference objective, which limits their ability to model the structured and sometimes conflicting nature of real-world human judgments that span multiple preference dimensions. In this work, we propose Listwise Direct Preference Optimization ($λ$-DPO), a unified framework that simultaneously improves supervision granularity and preference flexibility. Instead of collapsing multi-dimensional preference signals into a single ranking, $λ$-DPO constructs a mixture of listwise preference distributions weighted by a preference vector $λ$ on the probability simplex, enabling a single model to internalize a continuous spectrum of preference trade-offs. To further improve robustness, we introduce a performance-driven stochastic $λ$ scheduler that adaptively samples preference weights based on empirical downstream performance, explicitly mitigating the risks of misspecification inherent to static weighting schemes. We evaluate our method across multiple model families and scales on six widely used benchmarks. Experimental results show the consistent improvement against baselines.

偏好优化多维偏好模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。