让奖励模型学会识别不确定性的可靠预测。
Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
- 用概率输出建模人类偏好的随机性,捕捉内在不确定性。
- 集成多个模型差异,量化认知不确定性,识别不可靠判断。
- 不确定性低的预测更可靠,显著提升大模型生成质量。
奖励模型(RMs)对齐大语言模型(LLM)与人类期望至关重要。然而,现有模型难以捕捉人类偏好的随机性和不确定性,也无法评估奖励预测的可靠性。为此,我们提出不确定性感知奖励模型(URM)及其集成变体URME。URM通过概率价值头建模解耦的人类偏好属性分布,以捕获熵型不确定性;URME进一步通过集成内各模型间的差异量化认知不确定性,实现对不可靠评估的识别。实证表明,URM在RewardBench上表现优异,超越多个大规模基线模型。大量实验(包括best-of-n采样、迭代直接偏好优化和近端策略优化)显示,使用URM与URME可显著提升LLM生成质量。值得注意的是,不确定性较低的奖励预测更可靠,生成内容质量更高,对齐效果也更优。
原文摘要 · Abstract (English)
Reward models (RMs) are essential for aligning large language models (LLM) with human expectations. However, existing RMs struggle to capture the stochastic and uncertain nature of human preferences and fail to assess the reliability of reward predictions. To address these challenges, we introduce the Uncertainty-aware Reward Model (URM) and its ensemble variant, URME. URM employs a probabilistic value head to capture aleatoric uncertainty by modeling the distribution of disentangled human preference attributes. URME further quantifies epistemic uncertainty by examining discrepancies among individual URMs within the ensemble, enabling identification of unreliable evaluations. Our empirical evaluations demonstrate that URM achieves strong performance on RewardBench, outperforming competitive large-scale models. Additionally, extensive experiments, including best-of-n sampling (BoN), iterative direct preference optimization (iterative DPO), and proximal policy optimization (PPO), demonstrate that URM and URME significantly enhance LLMs' generation quality. Notably, reward predictions with lower uncertainty are far more reliable, demonstrate significantly higher quality, and result in substantially improved alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。