提出新评估框架,揭示主流对齐方法在多元偏好下存在严重性能失真。
Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
- 用社会选择理论建模用户偏好,引入'失真度'衡量对齐效果
- 发现RLHF和DPO在多元场景下平均效用损失可达β甚至无界
- 新方法纳什反馈学习可实现最优失真,鲁棒性强于现有技术
大语言模型在预训练后通过成对比较进行人类偏好对齐。当前主流方法(如基于PPO的RLHF和DPO)假设存在单一偏好模型,但实际部署中用户偏好多样。这导致这些方法是否能平均满足用户尚不明确——这是多元对齐的基本要求。本文基于社会选择理论,将用户比较建模为个体布拉德利-特雷西(BT)模型,提出对齐方法的失真度:最优可达平均效用与学习策略平均效用之比的最坏情况。该概念可清晰区分不同方法:纳什反馈学习(Nash LfH)在任意效用分布、比较对分布及允许的参考策略KL散度下,达到最小最大失真(1/2 + o(1))·β(β为BT温度)。而RLHF和DPO即使无KL约束,失真也≥(1 - o(1))·β;在完整设置下,失真可达e^{Ω(β)}甚至无界,取决于比较对采样方式。
原文摘要 · Abstract (English)
After pre-training, large language models are aligned with human preferences based on pairwise comparisons. State-of-the-art alignment methods (such as PPO-based RLHF and DPO) are built on the assumption of aligning with a single preference model, despite being deployed in settings where users have diverse preferences. As a result, it is not even clear that these alignment methods produce models that satisfy users on average -- a minimal requirement for pluralistic alignment. Drawing on social choice theory and modeling users' comparisons through individual Bradley-Terry (BT) models, we introduce an alignment method's distortion: the worst-case ratio between the optimal achievable average utility, and the average utility of the learned policy. The notion of distortion helps draw sharp distinctions between alignment methods: Nash Learning from Human Feedback achieves the minimax optimal distortion of $(\frac{1}{2} + o(1)) \cdot β$ (for the BT temperature $β$), robustly across utility distributions, distributions of comparison pairs, and permissible KL divergences from the reference policy. RLHF and DPO, by contrast, suffer $\geq (1 - o(1)) \cdot β$ distortion already without a KL constraint, and $e^{Ω(β)}$ or even unbounded distortion in the full setting, depending on how comparison pairs are sampled.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。