提出统一框架RewardUQ,评估奖励模型的不确定性量化效果。
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
- 构建统一评估框架,对比多种不确定性量化方法。
- 发现模型规模与初始化对性能影响最大,优于现有设计。
- 开源工具包,助力后续研究与实际应用部署。
奖励模型在对齐大语言模型与人类偏好中起核心作用。然而,现有方法多依赖点估计奖励,忽略了因人类反馈有限而产生的认知不确定性。近期研究表明,量化不确定性可通过主动学习降低人工标注成本,并缓解大语言模型后训练中的奖励过优化问题。但目前不确定性感知的奖励模型缺乏系统比较,理解不足。本文提出统一框架RewardUQ,系统评估奖励模型的不确定性量化方法。我们在标准指标上对比常见方法的准确率与校准性,并提出一种结合两者的新排序策略以简化比较。实验表明,模型规模与初始化对性能影响最为显著,多数前期工作本可受益于其他设计选择。为促进新方法开发与下游应用,我们开源了该框架的Python包,代码见https://github.com/lasgroup/rewarduq。
原文摘要 · Abstract (English)
Reward models are central to aligning large language models (LLMs) with human preferences. Yet most approaches rely on pointwise reward estimates that overlook the epistemic uncertainty in reward models arising from limited human feedback. Recent work suggests that quantifying this uncertainty can reduce the costs of human annotation via uncertainty-guided active learning and mitigate reward overoptimization in LLM post-training. However, uncertainty-aware reward models have so far been adopted without thorough comparison, leaving them poorly understood. This work introduces a unified framework, RewardUQ, to systematically evaluate uncertainty quantification for reward models. We compare common methods along standard metrics measuring accuracy and calibration, and we propose a new ranking strategy incorporating both dimensions for a simplified comparison. Our experimental results suggest that model size and initialization have the most meaningful impact on performance, and most prior work could have benefited from alternative design choices. To foster the development and evaluation of new methods and aid the deployment in downstream applications, we release our open-source framework as a Python package. Our code is available at https://github.com/lasgroup/rewarduq.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。