arXiv:2508.13189stat.MLcs.AI2025-08

揭示偏好模型与生存分析模型的深层关联

Preference Models assume Proportional Hazards of Utilities

  • 将普莱克特-卢斯模型与考克斯比例风险模型关联
  • 发现偏好假设隐含比例风险特性
  • 为对齐模型提供理论依据,适合研究者阅读

从人类标注数据中估计偏好的方法通常涉及对选择排序列表的分布建模,如普莱克特-卢斯(Plackett-Luce)模型。现代人工智能对齐工具,如奖励建模和直接偏好优化,均基于普莱克特-卢斯模型的统计假设。本文将普莱克特-卢斯模型与另一经典统计模型——考克斯比例风险模型联系起来,探讨二者之间的内在关联,并揭示这一联系带来的理论启示。

原文摘要 · Abstract (English)

Approaches for estimating preferences from human annotated data typically involves inducing a distribution over a ranked list of choices such as the Plackett-Luce model. Indeed, modern AI alignment tools such as Reward Modelling and Direct Preference Optimization are based on the statistical assumptions posed by the Plackett-Luce model. In this paper, I will connect the Plackett-Luce model to another classical and well known statistical model, the Cox Proportional Hazards model and attempt to shed some light on the implications of the connection therein.

偏好建模统计推断对齐理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。