通过平滑版最佳选择提升生成模型推理对齐效果
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
- 用软化版本的Best-of-N分析生成模型对齐机制
- 平滑策略能有效降低奖励过优化问题,尤其在代理奖励差时
- 适合研究生成模型对齐与奖励建模的科研人员
生成模型推理阶段对齐的一种简单而有效的方法是最佳选择N(BoN),即从参考策略中采样N个输出,利用代理奖励模型评估并选取得分最高的一个。尽管先前工作认为BoN在奖励与KL权衡上近乎最优,但其效果高度依赖于代理奖励模型的质量。为此,本文通过一种平滑版本——软最佳选择N(SBoN)进行研究,并建立理论框架加以补充。我们分析了BoN的缩放行为,给出了SBoN策略与参考策略之间KL散度的边界,揭示了性能随样本数变化的规律。同时研究了遗憾差距,即最优策略与SBoN策略在真实期望奖励上的差距。理论与实证结果表明,平滑机制有助于缓解奖励过优化问题,尤其在代理奖励质量较低时表现更优。
原文摘要 · Abstract (English)
A simple yet effective method for inference-time alignment of generative models is Best-of-$N$ (BoN), where $N$ outcomes are sampled from a reference policy, evaluated using a proxy reward model, and the highest-scoring one is selected. While prior work argues that BoN is almost optimal in reward vs KL tradeoffs, the effectiveness of BoN depends critically on the quality of the proxy reward model used for selection. For this purpose, we study BoN through a smooth version known as Soft Best-of-N (SBoN) and develop a theoretical framework to address this gap. We analyze the scaling behaviour of BoN by providing bounds on the KL divergence between the SBoN policy and the reference policy, offering insights into how performance varies with the number of samples. We also study the regret gap, i.e., the gap between the expected true reward under the optimal policy and the SBoN policy. Our theoretical and empirical findings show that smoothing helps SBoN mitigate reward overoptimization, especially when the quality of the proxy reward is low.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。