提出新基准RewardMATH,更真实评估数学推理奖励模型的鲁棒性。
Evaluating Robustness of Reward Models for Mathematical Reasoning
- 设计多样本对比机制,避免单次比较导致的偏差
- RewardMATH得分与优化策略结果强相关,能有效检测奖励过优
- 适合关注强化学习对齐与奖励模型可靠性的研究者
奖励模型在基于人类反馈的强化学习(RLHF)系统中至关重要,用于对齐模型行为与人类偏好。尤其在数学领域,已有大量研究使用奖励模型来提升推理能力。近期,RewardBench被提出以理解奖励模型的行为,但我们发现其数学子集在选择与拒绝完成项之间存在表示差异,且仅依赖单一比较,可能因仅观察孤立案例而产生不可靠结果,无法准确反映奖励模型的鲁棒性,导致对其性能的误解,甚至引发奖励劫持。为此,本文提出一种新的可靠评估设计,并构建了RewardMATH基准,有效体现数学推理任务中奖励模型的鲁棒性。实验表明,RewardMATH得分与优化策略结果高度相关,可有效估计奖励过优现象,而现有基准几乎无相关性。结果证明该设计显著提升了评估可靠性,更好地刻画了奖励模型的鲁棒性。代码与数据已公开。
原文摘要 · Abstract (English)
Reward models are key in reinforcement learning from human feedback (RLHF) systems, aligning the model behavior with human preferences. Particularly in the math domain, there have been plenty of studies using reward models to align policies for improving reasoning capabilities. Recently, as the importance of reward models has been emphasized, RewardBench is proposed to understand their behavior. However, we figure out that the math subset of RewardBench has different representations between chosen and rejected completions, and relies on a single comparison, which may lead to unreliable results as it only see an isolated case. Therefore, it fails to accurately present the robustness of reward models, leading to a misunderstanding of its performance and potentially resulting in reward hacking. In this work, we introduce a new design for reliable evaluation of reward models, and to validate this, we construct RewardMATH, a benchmark that effectively represents the robustness of reward models in mathematical reasoning tasks. We demonstrate that the scores on RewardMATH strongly correlate with the results of optimized policy and effectively estimate reward overoptimization, whereas the existing benchmark shows almost no correlation. The results underscore the potential of our design to enhance the reliability of evaluation, and represent the robustness of reward model. We make our code and data publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。