提出用三选一偏好数据改进人类偏好建模,更精准捕捉复杂偏好关系。
Learning Correlated Reward Models: Statistical Barriers and Opportunities
- 采用三选一偏好数据替代传统成对比较,避免独立无关选项假设
- 理论证明三选一数据可实现近最优统计与计算效率
- 在真实数据集上验证了个性化偏好建模效果提升
随机效用模型(RUM)是建模用户偏好的经典框架,在基于人类反馈的强化学习(RLHF)中至关重要。然而,许多方法依赖于无关选项独立性(IIA)假设,将所有人类偏好简化为单一效用函数,导致对人类偏好的粗略近似。而规避该假设的模型缺乏统计与计算保障。本文研究避免IIA假设的关联概率模型的统计与计算挑战。首先证明传统成对偏好数据根本无法学习相关性信息,解释了该设置下缺乏理论保证的原因。接着表明三选一偏好数据可有效克服此缺陷,并设计出具有近最优性能的统计与计算高效估计器。理论结果凸显高阶偏好数据在建模关联效用中的优势,支持更精细的人类偏好表达。最后在多个真实数据集上验证了理论保障,展示了偏好建模个性化能力的显著提升。
原文摘要 · Abstract (English)
Random Utility Models (RUMs) are a classical framework for modeling user preferences and play a key role in reward modeling for Reinforcement Learning from Human Feedback (RLHF). However, a crucial shortcoming of many of these techniques is the Independence of Irrelevant Alternatives (IIA) assumption, which collapses \emph{all} human preferences to a universal underlying utility function, yielding a coarse approximation of the range of human preferences. On the other hand, statistical and computational guarantees for models avoiding this assumption are scarce. In this paper, we investigate the statistical and computational challenges of learning a \emph{correlated} probit model, a fundamental RUM that avoids the IIA assumption. First, we establish that the classical data collection paradigm of pairwise preference data is \emph{fundamentally insufficient} to learn correlational information, explaining the lack of statistical and computational guarantees in this setting. Next, we demonstrate that \emph{best-of-three} preference data provably overcomes these shortcomings, and devise a statistically and computationally efficient estimator with near-optimal performance. These results highlight the benefits of higher-order preference data in learning correlated utilities, allowing for more fine-grained modeling of human preferences. Finally, we validate these theoretical guarantees on several real-world datasets, demonstrating improved personalization of human preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。