arXiv:2512.03208stat.MLcs.LG2025-12被引 8

为异质人类反馈下的大模型奖励学习提供不确定性量化方法

Uncertainty Quantification for Large Language Model Reward Learning under Heterogeneous Human Feedback

  • 构建异质偏好框架,联合建模答案真实奖励与人类理性
  • 提出交替梯度下降算法求解,理论证明收敛性与渐近分布
  • 可生成置信区间,提升奖励比较与最佳选优策略的可靠性

我们研究用于对齐大语言模型(LLM)的奖励模型的估计与统计推断。大模型对齐的关键是基于人类反馈的强化学习(RLHF),其中人类对模型生成的回答进行成对比较,其偏好用于训练奖励模型。然而,人类反馈本质上具有异质性,给可靠奖励学习带来挑战。为此,我们采用异质偏好框架,联合建模回答的潜在奖励与人类理性。这导致一个复杂的双凸优化问题,我们通过交替梯度下降算法求解。我们建立了该估计器的理论保证,包括收敛性与渐近分布。这些结果使得奖励估计的置信区间得以构建。利用不确定性量化结果,我们实现了奖励的有效统计比较,并将不确定性纳入最佳选优-$N$(BoN)策略框架。大量模拟实验验证了方法的有效性,真实大模型数据的应用凸显了在奖励建模中考虑不确定性的实际价值。

原文摘要 · Abstract (English)

We study estimation and statistical inference for reward models used in aligning large language models (LLMs). A key component of LLM alignment is reinforcement learning from human feedback (RLHF), where humans compare pairs of model-generated answers and their preferences are used to train a reward model. However, human feedback is inherently heterogeneous, creating significant challenges for reliable reward learning. To address this, we adopt a heterogeneous preference framework that jointly models the latent reward of answers and human rationality. This leads to a challenging biconvex optimization problem, which we solve via an alternating gradient descent algorithm. We establish theoretical guarantees for the resulting estimator, including its convergence and asymptotic distribution. These results enable the construction of confidence intervals for reward estimates. Leveraging these uncertainty quantification results, we conduct valid statistical comparisons between rewards and incorporate uncertainty into the best-of-$N$ (BoN) policy framework. Extensive simulations demonstrate the effectiveness of our method, and applications to real LLM data highlight the practical value of accounting for uncertainty in reward modeling for LLM alignment.

奖励学习不确定性量化人类反馈大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。