奖励学习中,表示能力越强,越难给出一致评分。
The Representation-Rationalizability Tradeoff in Reward Learning

- 用嵌入表示区分响应,影响可比较的选项范围
- 表示越丰富,跨熵损失中不可一致性部分越大
- 适合关注奖励模型设计权衡的研究者
在人类反馈强化学习(RLHF)中,每个训练样本包含一个提示x和两个候选响应y、y′,标注者提供两者之间的成对偏好。学习目标是将这些异质的成对判断转化为单个标量奖励r(x,y),以衡量每个提示下响应的质量。经典社会选择理论表明,由于标注者样本异质性可能引发凝聚偏好循环(Condorcet cycles),因此无法存在一个标量奖励能始终一致地评估所有响应对。现有文献常假设每提示有固定有限的候选响应集,但现代流程先通过学习的表示ϕ(x,y)对响应打分,再经标量头输出奖励。此时嵌入ϕ决定了哪些响应被视为可区分的备选项,哪些比较可见于奖励模型。一旦嵌入成为问题的一部分,社会选择中的不可能性便转化为一种权衡:我们证明,任意基于ϕ构建的奖励的额外交叉熵损失可精确分解为两部分——表示项随ϕ丰富而减小,聚合项则因暴露更多无法一致排序的比较而增大。该结论同样适用于直接偏好优化(DPO),且联合训练嵌入与奖励无法保证达到这一权衡的最佳平衡点。在合成数据和真实偏好数据集上的实验验证了上述结果。
原文摘要 · Abstract (English)
In RLHF, each training example contains a prompt $x$ and two candidate responses $y,y'$, and annotators provide pairwise preferences between these responses. The learning problem is to convert these heterogeneous pairwise judgments into a single scalar reward $r(x,y)$ that measures response quality for each prompt. Classical social choice implies an impossibility because heterogeneous annotator samples can induce pooled preferences with Condorcet cycles, so no scalar reward can evaluate all compared response pairs consistently. A growing literature analyzes RLHF as a social-choice problem, but usually assumes a fixed finite set of alternatives, i.e., a pre-enumerated finite set of candidate responses for each prompt. Modern pipelines instead score responses through a learned representation $ϕ(x,y)$ before a scalar head, so $ϕ$ determines which responses are treated as distinguishable alternatives and which comparisons are visible to the reward model. Once this embedding is part of the problem, the impossibility results from social choice theory become a tradeoff. We show that the excess cross-entropy loss of any reward built on $ϕ$ decomposes exactly into a representational term, which a richer $ϕ$ shrinks, and an aggregation term, which a richer $ϕ$ enlarges by exposing more comparisons that no scalar can rank consistently. The same results extend to direct preference optimization (DPO), and jointly training the embedding and the reward cannot guarantee to recover the sweet spot of this tradeoff. Experiments on synthetic data and real preference datasets corroborate our results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。