新评测基准让奖励模型得分平均低20分,更贴近实际应用效果。
RewardBench 2: Advancing Reward Model Evaluation
- 构建全新多技能评测集,使用人工生成新提示增强严谨性
- 现有模型在新基准上平均得分下降约20分,与下游任务表现高度相关
- 适合评估奖励模型在推理和强化学习中的真实效能
奖励模型在语言模型后训练中用于捕捉偏好数据中的细微信号,并为指令遵循、推理、安全等领域的优化提供训练目标。社区已开始建立奖励模型评估的最佳实践,包括针对特定技能的测试基准以及与人类偏好一致性的测试。然而,评估进展并未反映在奖励模型下游任务中的实际效果——许多情况下,更简单的直接对齐算法表现更好。本文提出 RewardBench 2,一个全新的多技能奖励模型评测基准,旨在提供更具挑战性的数据以推动基于准确率的评估。相比 RewardBench 1,模型在 RewardBench 2 上平均得分降低约 20 分,且与下游任务性能高度相关。与其他基准不同,RewardBench 2 使用人工生成的新提示,而非来自下游评估的现成提示,有助于实现更严格的评估实践。本文详细描述了基准构建过程,并报告了现有模型的表现,同时量化了其在推理阶段的缩放算法(如 best-of-N 采样)和强化学习中的训练算法(如近端策略优化,PPO)中的下游表现相关性。
原文摘要 · Abstract (English)
Reward models are used throughout the post-training of language models to capture nuanced signals from preference data and provide a training target for optimization across instruction following, reasoning, safety, and more domains. The community has begun establishing best practices for evaluating reward models, from the development of benchmarks that test capabilities in specific skill areas to others that test agreement with human preferences. At the same time, progress in evaluation has not been mirrored by the effectiveness of reward models in downstream tasks -- simpler direct alignment algorithms are reported to work better in many cases. This paper introduces RewardBench 2, a new multi-skill reward modeling benchmark designed to bring new, challenging data for accuracy-based reward model evaluation -- models score about 20 points on average lower on RewardBench 2 compared to the first RewardBench -- while being highly correlated with downstream performance. Compared to most other benchmarks, RewardBench 2 sources new human prompts instead of existing prompts from downstream evaluations, facilitating more rigorous evaluation practices. In this paper, we describe our benchmark construction process and report how existing models perform on it, while quantifying how performance on the benchmark correlates with downstream use of the models in both inference-time scaling algorithms, like best-of-N sampling, and RLHF training algorithms like proximal policy optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。