arXiv:2504.14838cs.AI2025-04被引 1

提出新指标RETA,量化奖励模型可靠性。

Establishing Reliability Metrics for Reward Models in Large Language Models

  • 用顶级响应的平均质量评估奖励模型可靠性
  • 实验证明该指标稳定,可识别最优采样量化阶
  • 无需额外人工标注即可评估任意奖励模型

奖励模型(RM)在大语言模型优化中至关重要,例如通过人类反馈强化学习(RLHF)或拒绝采样。然而,现有奖励模型的可靠性存疑:高分输出未必符合真实人类偏好。当前缺乏可信的可靠性度量标准。为此,我们提出【可靠在η】(RETA)指标,通过评估奖励模型排名前η分位数的输出在权威评分下的平均质量,直接衡量其可靠性。基于RETA,我们构建了一个集成基准测试流程,使任何人都可在不增加人工标注成本的前提下评估自身奖励模型。大量实验表明,RETA指标具有优异稳定性,能有效评估多个公开及私有奖励模型的可靠性。当面对不可靠的奖励模型时,可借助RETA确定最佳响应采样分位数。

原文摘要 · Abstract (English)

The reward model (RM) that represents human preferences plays a crucial role in optimizing the outputs of large language models (LLMs), e.g., through reinforcement learning from human feedback (RLHF) or rejection sampling. However, a long challenge for RM is its uncertain reliability, i.e., LLM outputs with higher rewards may not align with actual human preferences. Currently, there is a lack of a convincing metric to quantify the reliability of RMs. To bridge this gap, we propose the \textit{\underline{R}eliable at \underline{$η$}} (RETA) metric, which directly measures the reliability of an RM by evaluating the average quality (scored by an oracle) of the top $η$ quantile responses assessed by an RM. On top of RETA, we present an integrated benchmarking pipeline that allows anyone to evaluate their own RM without incurring additional Oracle labeling costs. Extensive experimental studies demonstrate the superior stability of RETA metric, providing solid evaluations of the reliability of various publicly available and proprietary RMs. When dealing with an unreliable RM, we can use the RETA metric to identify the optimal quantile from which to select the responses.

奖励模型可靠性评估大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。