用聊天机器人对战评分校准奖励模型,让AI评价更公平可靠。
CHARM: Calibrating Reward Models With Chatbot Arena Scores
- 基于聊天机器人对战的Elo评分构建去偏偏好数据集
- 校准后模型在多个评测集上与人类偏好相关性提升
- 适合需要公平评价的RLHF训练场景
奖励模型(RMs)在基于人类反馈的强化学习中起关键作用,作为人类偏好的代理以对齐大语言模型。然而,它们存在多种偏差,可能导致奖励欺骗。本文发现一种模型偏好偏差:某些策略模型的回复会被系统性地赋予过高分数,造成不公平判断。为此,我们提出一种名为CHARM的校准方法,利用聊天机器人对战(Chatbot Arena)的Elo评分构建去偏偏好数据集,并调整奖励模型得分。我们在奖励模型基准和人类偏好对齐任务上进行了广泛实验。结果表明,校准后的奖励模型在RM-Bench和RewardBench的Chat-Hard领域评估准确率更高,与人类偏好相关性更强,生成的分数更贴近Elo排名,并提升了下游微调性能。这些结果证明CHARM是一种简单、有效且通用的可靠奖励模型构建方法。
原文摘要 · Abstract (English)
Reward models (RMs) play a crucial role in Reinforcement Learning from Human Feedback by serving as proxies for human preferences in aligning large language models. However, they suffer from various biases which could lead to reward hacking. In this paper, we identify a model preference bias in RMs, where they systematically assign disproportionately high scores to responses from certain policy models, leading to unfair judgments. To mitigate this bias, we propose a calibration method named CHatbot Arena calibrated Reward Modeling (CHARM) that leverages Elo scores from the Chatbot Arena to construct debiased preference datasets and adjust reward model scoring. We conduct extensive experiments on reward model benchmarks and human preference alignment. Results demonstrate that our calibrated RMs achieve improved evaluation accuracy on RM-Bench and the Chat-Hard domain of RewardBench, exhibit a stronger correlation with human preferences by producing scores more closely aligned with Elo rankings and improve downstream post-training performance. These results demonstrate that CHARM provides a simple, effective, and broadly applicable approach to building more reliable and fair reward models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。