用彩票式投票机制替代传统对齐方法,让AI更尊重多数人意愿。
Jackpot! Alignment as a Maximal Lottery
- 用最大彩票规则替代RLHF,基于概率选择最优回应
- 实验表明新方法能更好处理偏好矛盾与无关选项干扰
- 适合需要公平对齐、重视多数意见的AI系统设计
强化学习人类反馈(RLHF)是当前对齐大语言模型与人类价值观的标准方法,但常无法满足直观上理想的属性,如尊重多数人偏好。为此,我们提出以概率性社会选择规则——最大彩票(maximal lotteries)取代RLHF。我们证明,一类对齐技术(如纳什人类反馈学习NLHF及其变体)近似于最大彩票结果,因而继承其优良性质。实验验证,新方法在处理偏好数据时更具鲁棒性:能支持多数人偏好,合理应对偏好非传递性,并对无关选项不敏感。这使得系统更能体现人类价值观与意图。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF), the standard for aligning Large Language Models (LLMs) with human values, is known to fail to satisfy properties that are intuitively desirable, such as respecting the preferences of the majority \cite{ge2024axioms}. To overcome these issues, we propose the use of a probabilistic Social Choice rule called \emph{maximal lotteries} as a replacement for RLHF. We show that a family of alignment techniques, namely Nash Learning from Human Feedback (NLHF) \cite{munos2023nash} and variants, approximate maximal lottery outcomes and thus inherit its beneficial properties. We confirm experimentally that our proposed methodology handles situations that arise when working with preferences more robustly than standard RLHF, including supporting the preferences of the majority, providing principled ways of handling non-transitivities in the preference data, and robustness to irrelevant alternatives. This results in systems that better incorporate human values and respect human intentions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。