研究如何从最佳选一数据中更有效学习奖励模型,揭示了样本数量与生成分布的设计原则。
Reward Learning from Best-of-$N$ Preference Data: Targets, Tradeoffs, and Design Principles

- 基于贝叶斯-特里模型,推导出最佳选一数据的精确奖励目标函数。
- 发现大样本量提升置信度但降低连通性,需权衡生成瓶颈与标注瓶颈。
- 提出调整基础生成分布以聚焦关键对比,实验验证设计有效性。
最佳选一(Best-of-$N$)采样广泛用于构建成对偏好数据:从基础分布中抽取 $N$ 个候选,选择最优者与被拒响应配对。尽管应用广泛,其对奖励学习的影响及 $N$ 和基础分布的选择仍不明确。本文将近期关于偏好数据诱导条件分布的分析应用于最佳选一场景。对于独立参考变体,推导出奖励目标的闭式表达,明确依赖于 $N$ 与基础分布,并保持潜在奖励排序。在实际常用变体(最佳对随机、最佳对最差)中,选中与拒绝响应共享同一候选集,导致严格贝叶斯-特里可表示性通常失效;但有界类最小化器随 $N$ 增大趋近参考目标。已知边际与连通性决定成对偏好学习的样本效率,而最佳选一通过 $N$ 将二者以相反方向耦合:增大 $N$ 可扩大成对边际,但降低连通性。由此得出两条设计原则:当偏好标签为瓶颈时用较大 $N$,生成为瓶颈时用较小 $N$;并调整基础分布,使质量集中在测试时最需比较的响应之间。合成与真实偏好数据实验支持样本量与分布形状的预测影响。
原文摘要 · Abstract (English)
Best-of-$N$ sampling is widely used to construct pairwise preference data: $N$ candidates are drawn from a base distribution, and the best is paired with a rejected response. Despite its widespread use, what Bradley--Terry (BT) reward learning extracts from such data, and how to choose $N$ and the base distribution, remain unclear. We specialize a recent analysis of preference data via its induced conditional distribution to Best-of-$N$. For independent-reference variants, we derive closed-form reward targets as explicit functions of $N$ and the base distribution, and show that they preserve the latent reward ranking. For the practical Best-vs-Random and Best-vs-Worst variants, chosen and rejected responses are coupled through the same candidate set, so exact BT representability generally fails; nevertheless, bounded-class minimizers approach the reference targets as $N$ grows. Although margin and connectivity are known to govern sample efficiency in pairwise preference learning, Best-of-$N$ couples them through $N$ in opposing directions: larger $N$ widens pairwise margins but reduces connectivity. This trade-off yields two design principles: use larger $N$ when preference labels are the bottleneck, smaller $N$ when generation is the bottleneck; and shape the base distribution to place mass between the responses whose comparison matters most at test time. Experiments on synthetic and real preference data support the predicted dependence on sample size and base-distribution shape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。