用锚点标签解决奖励模型方差识别难题,提升多样化偏好建模效果。
Variance-aware Reward Modeling with Anchor Guidance

- 引入两个粗粒度锚点标签,破解方差模型不可识别问题。
- 在四个真实数据集上,奖励建模与强化学习效果均显著提升。
- 适合处理人类偏好多样、需建模不确定性的场景。
标准的Bradley-Terry(BT)奖励模型在人类偏好多样化时表现受限。尽管软偏好标签能保留分歧信息,但BT只能通过缩小奖励间隔来表达,导致性能下降。高斯奖励模型通过联合预测奖励均值和方差提供替代方案,但仅凭成对偏好存在根本性不可识别性。本文提出锚点引导的方差感知奖励建模框架,通过在偏好数据中加入两个粗粒度响应级锚点标签,解决了该不可识别性问题。基于此,我们证明两个锚点足以实现识别,设计了联合训练目标,并建立了奖励均值与方差函数的非渐近收敛速率。在模拟实验及四个真实世界分歧偏好数据集上的实验表明,该方法在奖励建模性能和下游强化学习人类反馈(RLHF)任务中均持续领先,包括PPO训练和最优选择-$N$策略。
原文摘要 · Abstract (English)
Standard Bradley--Terry (BT) reward models are limited when human preferences are pluralistic. Although soft preference labels preserve disagreement information, BT can only express it by shrinking reward margins. Gaussian reward models provide an alternative by jointly predicting a reward mean and a reward variance, but suffer from a fundamental non-identifiability from pairwise preferences alone. We propose Anchor-guided Variance-aware Reward Modeling, a framework that resolves this non-identifiability by augmenting preference data with two coarse response-level anchor labels. Building on this, we prove that two anchors are sufficient for identification, develop a joint training objective and establish a non-asymptotic convergence rate for both the estimated reward mean and variance functions. Across simulation studies and four real-world diverging-preference datasets, our method consistently improves reward modeling performance and downstream RLHF, including PPO training and best-of-$N$ selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。