解决大模型对齐中用户偏好差异问题,用三元比较提升公平性与个性化。
Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
- 引入三元比较替代二元反馈,确保用户偏好可识别。
- 提出基于期望最大化的新算法,自动发现不同用户类型并分别训练模型。
- 设计最小最大后悔公平策略,保证所有用户群体表现均衡。
强化学习从人类反馈(RLHF)已成为对齐大语言模型与人类价值观的核心方法,通常先通过偏好数据训练奖励模型,再用强化学习更新模型。近期的直接偏好优化(DPO)简化了这一流程,直接在偏好上进行优化。然而,这些方法常假设标注者偏好一致,且依赖二元比较,忽略了人类评估者的多样性及成对反馈的局限性。本文将RLHF中的偏好学习与计量经济学文献关联,证明仅用有限用户数据下的二元比较无法识别潜在用户偏好,而三元及以上响应的完整或不完整排序可确保可识别性。我们提出了一种适用于DPO的期望最大化算法,能发现隐含的标注者类型并相应地训练混合语言模型。此外,提出一种基于最小最大后悔公平准则的聚合算法,生成具有公平性能保障的单一生成策略。这些贡献共同构建了一个面向多样化用户的生成模型对齐理论与算法框架。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) has become central to aligning large language models with human values, typically by first learning a reward model from preference data which is then used to update the model with reinforcement learning. Recent alternatives such as Direct Preference Optimization (DPO) simplify this pipeline by directly optimizing on preferences. However, both approaches often assume uniform annotator preferences and rely on binary comparisons, overlooking two key limitations: the diversity of human evaluators and the limitations of pairwise feedback. In this work, we address both these issues. First, we connect preference learning in RLHF with the econometrics literature and show that binary comparisons are insufficient for identifying latent user preferences from finite user data and infinite users, while (even incomplete) rankings over three or more responses ensure identifiability. Second, we introduce methods to incorporate heterogeneous preferences into alignment algorithms. We develop an Expectation-Maximization adaptation of DPO that discovers latent annotator types and trains a mixture of LLMs accordingly. Then we propose an aggregation algorithm using a min-max regret fairness criterion to produce a single generative policy with equitable performance guarantees. Together, these contributions establish a theoretical and algorithmic framework for fairness and personalization for diverse users in generative model alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。