重新评估了BoN在推理对齐中的有效性,发现其在实际指标下表现最优。
Revisiting the (Sub)Optimality of Best-of-N for Inference-Time Alignment
- 基于实际常用的胜率指标分析BoN,而非理论假设的期望真实奖励
- 在合理条件下,调优后的BoN在胜率上达到统计与计算最优
- 提出新变体可彻底消除奖励劫持,适合注重鲁棒性的实际部署
Best-of-N(BoN)是一种广泛使用的语言模型推理时对齐方法,即从参考模型中采样N个候选回复,选择奖励模型预测得分最高的一个。尽管应用广泛,已有理论研究指出其在统计上次优且易受奖励劫持影响。本文在更贴近实践的假设下重新审视该问题:不同于以往关注期望真实奖励(在许多实际场景中无意义),我们分析推理对齐对胜率(win-rate)的影响,该指标更贴合奖励模型训练与评估方式。结果表明,在参考模型和奖励模型质量满足最低条件时,经过适当调优的BoN在实现高胜率方面兼具计算与统计最优性,部分解释了其广泛应用的成功。由于BoN仍可能受奖励劫持影响,我们提出一种简单且实用的改进变体,理论上可完全消除奖励劫持同时保持最优统计性能。最后,我们证明先前方法在胜率目标下必然次优,强调分析推理对齐方法时需选择恰当目标。
原文摘要 · Abstract (English)
Best-of-N (BoN) sampling is a widely used inference-time alignment method for language models, whereby N candidate responses are sampled from a reference model and the one with the highest predicted reward according to a learned reward model is selected. Despite its widespread practical use, recent theoretical work has suggested that it is statistically suboptimal and vulnerable to reward hacking, the process by which models exploit weaknesses in the learned reward model to achieve high estimated reward without genuinely improving performance. We revisit this question under assumptions that more closely reflect practice than that of prior work. In particular, in contradistinction to earlier analyses that focused on expected true reward, which may not be meaningful in many practical settings, we investigate how inference-time alignment affects the win-rate, a pairwise comparison-based metric more closely aligned with how reward models are trained and evaluated in practice. We demonstrate that, under minimal conditions on the quality of the reference model and learned reward model, properly tuned BoN is both computationally and statistically optimal in achieving high win-rate, partially explaining its widespread practical success. Because BoN remains susceptible to reward-hacking in this setting, we propose a simple and practical variant that provably eliminates reward-hacking while maintaining optimal statistical performance. Finally, we show that prior approaches are provably suboptimal when considering win-rate, highlighting the importance of choosing appropriate objectives when analyzing inference-time alignment methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。