提出可调控的投票式奖励聚合框架,让RLHF更透明地选择偏好合并方式。
Differential Voting: Loss Functions For Axiomatically Diverse Aggregation of Heterogeneous Preferences
- 设计可微分损失函数,实现不同经典投票规则的精确建模
- 证明各损失对应的社会选择规则在极限下保持一致
- 揭示损失几何对公平性、稳定性等规范属性的影响
从人类反馈中进行强化学习(RLHF)隐式地将异质的人类偏好聚合为单一效用函数,实际上参与者的真实效用是多样的。因此,RLHF可被视为一种投票机制,其聚合方式由损失函数定义。尽管阿罗不可能定理表明不同机制满足不同的理想公理集,但现有方法大多依赖单一聚合原则,通常是对应于博达计票的布拉德利-特雷西-卢斯(BTL)模型。这限制了学习到的奖励函数的公理性质,并掩盖了优化中的规范假设。本文提出差分投票(Differential Voting)框架,构建实例级可微损失函数,其全局最优解严格对应于经典投票规则。我们开发了基于多数决(BTL)、科佩兰德和肯尼规则的可微代理,形式化分析其校准性、梯度场及平滑参数趋近零时的极限行为。对每种损失,我们建立与相应社会选择规则的一致性,并刻画其所满足或违背的公理。分析表明,损失几何设计(如边界敏感性与集中度)直接决定规范聚合行为。差分投票使偏好聚合成为可显式控制的设计选择,支持在公理保证与优化稳定性之间进行合理权衡。代码已开源。
原文摘要 · Abstract (English)
Reinforcement learning from human feedback (RLHF) implicitly aggregates heterogeneous human preferences into a single utility function, even though the underlying utilities of the participants are in practice diverse. Hence, RLHF can be viewed as a form of voting, where the aggregation mechanism is defined by the loss function. Although Arrow's Impossibility Theorem suggests that different mechanisms satisfy different sets of desirable axioms, most existing methods rely on a single aggregation principle, typically the Bradley-Terry-Luce (BTL) model, which corresponds to Borda count voting. This restricts the axiomatic properties of the learned reward and obscures the normative assumptions embedded in optimization. In this work, we introduce Differential Voting, a unifying framework that constructs instance-wise, differentiable loss functions whose population-level optima provably correspond to distinct classical voting rules. We develop differentiable surrogates for majority-based aggregation (BTL), Copeland, and Kemeny rules, and formally analyze their calibration properties, gradient fields, and limiting behavior as smoothing parameters vanish. For each loss, we establish consistency with the corresponding social choice rule and characterize the axioms it satisfies or violates. Our analysis shows how design choices in loss geometry-such as margin sensitivity and boundary concentration-directly translate into normative aggregation behavior. Differential Voting makes preference aggregation an explicit and controllable design choice in RLHF, enabling principled trade-offs between axiomatic guarantees and optimization stability. Code to reproduce our experiments is open-sourced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。