提出新评估框架,用过优化程度更真实衡量奖励模型能力。
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
- 从奖励过优化视角设计评估方法,关注响应差异与学习信号动态。
- 需多轮对比不同模型生成的响应,才能准确反映奖励模型性能。
- 适合研究RLHF对齐机制或构建可靠评测基准的研究者参考。
奖励模型(RMs)在人类反馈强化学习(RLHF)中至关重要,用于对齐模型行为与人类偏好。然而现有奖励模型评测基准与优化策略表现的相关性较弱,表明其无法准确评估奖励模型的真实能力。为此,本文从奖励过优化现象出发,探索多种评估设计。结果揭示三点关键:(i)应尽量减少选择与拒绝响应间的非正确性差异;(ii)评估需在广泛响应组合中进行多轮比较;(iii)因奖励模型需处理多样化表达,响应应来自多种模型。但我们也发现,与过优化程度高度相关反而导致与某些下游任务性能相关性下降。因此,过优化程度应作为评估工具而非最终目标。
原文摘要 · Abstract (English)
Reward models (RMs) play a crucial role in reinforcement learning from human feedback (RLHF), aligning model behavior with human preferences. However, existing benchmarks for reward models show a weak correlation with the performance of optimized policies, suggesting that they fail to accurately assess the true capabilities of RMs. To bridge this gap, we explore several evaluation designs through the lens of reward overoptimization\textemdash a phenomenon that captures both how well the reward model aligns with human preferences and the dynamics of the learning signal it provides to the policy. The results highlight three key findings on how to construct a reliable benchmark: (i) it is important to minimize differences between chosen and rejected responses beyond correctness, (ii) evaluating reward models requires multiple comparisons across a wide range of chosen and rejected responses, and (iii) given that reward models encounter responses with diverse representations, responses should be sourced from a variety of models. However, we also observe that a extremely high correlation with degree of overoptimization leads to comparatively lower correlation with certain downstream performance. Thus, when designing a benchmark, it is desirable to use the degree of overoptimization as a useful tool, rather than the end goal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。