解释了为何违背理论的RLHF仍能实用,并提出改进方法。
Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory
- 在现实偏好下,RLHF满足关键一致性条件。
- 微调目标可让其在任意偏好下保持一致性。
- 提出新标准,指导未来对齐方法设计。
尽管强化学习从人类反馈(RLHF)在实践中表现优异,但其违反了社会选择理论中的几乎所有基础公理,如多数一致性、成对多数一致性和康多塞一致性。这引发根本性疑问:为何它在实践中的表现如此出色?本文证明,在对偏好分布施加温和且符合实际的假设下,RLHF实际上满足成对多数一致性和康多塞一致性。这些假设在真实对齐任务中常成立,为其实用性能提供了理论解释。此外,我们提出对奖励建模目标进行微小修改,即可在一般偏好分布下保证成对多数或康多塞一致性,从而提升对齐效果。最后,我们超越传统经济与社会选择理论的公理框架,引入三个新对齐标准——偏好匹配、偏好等价和群体偏好匹配——更贴合学习响应分布的目标。结果显示,虽然RLHF满足前两项,但不满足第三项。文章最后讨论如何设计未来对齐方法以同时满足全部三项。
原文摘要 · Abstract (English)
Despite its empirical success, Reinforcement Learning from Human Feedback (RLHF) has been shown to violate almost all the fundamental axioms in social choice theory -- such as majority consistency, pairwise majority consistency, and Condorcet consistency. This raises a foundational question: why does RLHF perform so well in practice if it fails these seemingly essential properties? In this paper, we resolve this paradox by showing that under mild and empirically plausible assumptions on the preference profile, RLHF does satisfy pairwise majority and Condorcet consistency. These assumptions are frequently satisfied in real-world alignment tasks, offering a theoretical explanation for RLHF's strong practical performance. Furthermore, we show that a slight modification to the reward modeling objective can ensure pairwise majority or Condorcet consistency even under general preference profiles, thereby improving the alignment process. Finally, we go beyond classical axioms in economic and social choice theory and introduce new alignment criteria -- preference matching, preference equivalence, and group preference matching -- that better reflect the goal of learning distributions over responses. We show that while RLHF satisfies the first two properties, it fails to satisfy the third. We conclude by discussing how future alignment methods may be designed to satisfy all three.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。