构建全面的奖励模型评估基准,揭示现有方法的局限性。
RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
- 覆盖49种真实场景,采用成对与最佳N选一评估方式
- 发现主流奖励模型在泛化能力上的缺陷,验证其有效性
- 适合研究大模型对齐与奖励模型评估的学者使用
奖励模型(RMs)用于引导大语言模型(LLMs)的行为符合人类偏好。评估RMs是实现更好对齐的关键。然而,当前的评估可能无法准确反映其对齐性能,原因在于评估数据分布有限,且评估方法与对齐目标关联不紧密。为此,我们提出RMB,一个涵盖超过49种真实场景的综合性奖励模型评估基准,包含成对与Best-of-N(BoN)评估,以更准确反映RMs在引导对齐优化中的效果。我们证明了该基准与下游对齐任务性能之间存在正相关关系。基于此基准,我们对前沿奖励模型进行了广泛分析,揭示了以往基准未能发现的泛化缺陷,并突显了生成式奖励模型的潜力。此外,我们探讨了奖励模型中的开放问题,具体考察了多数投票在评估中的有效性,以及生成式奖励模型的影响因素,包括评估标准和指令方法的影响。评估代码与数据集已开源于https://github.com/Zhou-Zoey/RMB-Reward-Model-Benchmark。
原文摘要 · Abstract (English)
Reward models (RMs) guide the alignment of large language models (LLMs), steering them toward behaviors preferred by humans. Evaluating RMs is the key to better aligning LLMs. However, the current evaluation of RMs may not directly correspond to their alignment performance due to the limited distribution of evaluation data and evaluation methods that are not closely related to alignment objectives. To address these limitations, we propose RMB, a comprehensive RM benchmark that covers over 49 real-world scenarios and includes both pairwise and Best-of-N (BoN) evaluations to better reflect the effectiveness of RMs in guiding alignment optimization. We demonstrate a positive correlation between our benchmark and the downstream alignment task performance. Based on our benchmark, we conduct extensive analysis on the state-of-the-art RMs, revealing their generalization defects that were not discovered by previous benchmarks, and highlighting the potential of generative RMs. Furthermore, we delve into open questions in reward models, specifically examining the effectiveness of majority voting for the evaluation of reward models and analyzing the impact factors of generative RMs, including the influence of evaluation criteria and instructing methods. Our evaluation code and datasets are available at https://github.com/Zhou-Zoey/RMB-Reward-Model-Benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。