揭示大模型对齐人类偏好的统计极限,发现奖励方法有根本性缺陷,但非奖励方法可保留少数偏好。
Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium
- 用奖赏模型对齐需无康多塞循环,而该循环在人类偏好中几乎必然存在。
- 基于纳什学习的非奖赏对齐方法可让模型采用混合策略,避免单一输出。
- 在通用偏好模型下,保留少数偏好几乎总是可能,无需额外正则化。
将大语言模型(LLMs)与多样化的人类偏好对齐对决策部署中的公平性和可靠性至关重要。本文探究了此类对齐的根本统计限制,聚焦于人类偏好的概率表示及对齐后模型中多样性偏好的保留。我们首先证明:仅当模型生成响应间的偏好无康多塞循环时,才能用奖赏模型表示人类偏好。进一步,在广义的概率偏好模型——卢斯模型(Luce model)下,康多塞循环的存在概率以指数速度趋近于1,表明基于奖赏的方法(如从人类反馈中强化学习)无法实现完全对齐。接着,我们研究在非奖赏方法(如从人类反馈中纳什学习)下,模型在极限状态下采用混合策略(不固定为单一输出)的条件。我们确定了必要且充分条件:不存在被多数人始终偏好的响应。作为利好,该条件在卢斯模型下以高概率成立,表明在对齐过程中无需显式正则化即可保留少数偏好。
原文摘要 · Abstract (English)
Aligning large language models (LLMs) with diverse human preferences is critical for ensuring fairness and informed outcomes when deploying these models for decision-making. In this paper, we seek to uncover fundamental statistical limits concerning aligning LLMs with human preferences, with a focus on the probabilistic representation of human preferences and the preservation of diverse preferences in aligned LLMs. We first show that human preferences can be represented by a reward model if and only if the preference among LLM-generated responses is free of any Condorcet cycle. Moreover, we prove that Condorcet cycles exist with probability converging to one exponentially fast under a general probabilistic preference model called the Luce model, thereby demonstrating the impossibility of fully aligning human preferences using reward-based approaches such as reinforcement learning from human feedback. Next, we explore the conditions under which LLMs would employ mixed strategies -- meaning they do not collapse to a single response -- when aligned in the limit using a non-reward-based approach, such as Nash learning from human feedback. We identify a necessary and sufficient condition for mixed strategies: the absence of a response that is preferred over all others by a majority. As a blessing, we prove that this condition holds with high probability under the Luce model, thereby highlighting the statistical possibility of preserving minority preferences without explicit regularization in aligning LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。