arXiv:2602.08819cs.LGcs.CL2026-02

让奖励模型在测试时通过示例动态调整偏好,提升对复杂人类偏好的适应能力。

Bayesian Preference Learning for Test-Time Steerable Reward Models

  • 用贝叶斯推断和上下文示例实现测试时可调节的奖励建模
  • 多示例下RM-Bench准确率从60.5提升至70.8,校准误差更低
  • 适合需要动态对齐人类偏好的强化学习场景,如数学推理与道德困境

奖励模型在通过强化学习对齐语言模型与人类偏好中起核心作用。随着强化学习应用于可验证奖励和多目标对齐等场景,奖励模型需刻画更复杂、多维度的偏好分布。然而,传统分类器型奖励模型训练后即固定,难以在测试时灵活调整。本文提出变分上下文奖励建模(ICRM),一种新型贝叶斯奖励建模方法,可通过上下文偏好示例实现测试时可调性。ICRM将奖励建模视为在布拉德利-特瑞模型下对潜在偏好概率的摊销变分推断,采用共轭贝塔先验。我们证明ICRM可在单目标与多目标设置下适应未见偏好分布。示例越多,其在RM-Bench上的准确率从60.5提升至70.8,道德困境偏好校准误差低于生成式裁判,且在冲突偏好下扩展了可达帕累托前沿。进一步实验表明,ICRM能有效编码可验证奖励,在数学推理任务中优于传统奖励模型。理论分析显示,该变分目标存在有限置信度下的全局内点最优解,且KL正则化可缓解奖励过优化问题。

原文摘要 · Abstract (English)

Reward models are central to aligning language models with human preferences via reinforcement learning (RL). As RL is increasingly applied to settings such as verifiable rewards and multi-objective alignment, RMs are expected to encode more complex and multifaceted preference distributions. However, classifier RMs remain static once trained, limiting their adaptability at test time. We propose Variational In-Context Reward Modeling (ICRM), a novel Bayesian reward modeling objective that enables test-time steerability via in-context preference demonstrations. ICRM casts reward modeling as amortized variational inference over a latent preference probability under the Bradley-Terry model using a conjugate Beta prior. We show that ICRM adapts to unseen preference distributions at test time for both single and multi-objective settings. With more demonstrations, ICRM improves RM-Bench accuracy from 60.5 to 70.8, achieves lower calibration error than a generative judge on moral dilemma preferences, and expands the attainable Pareto frontier under conflicting preferences. We further study the practical applicability of ICRM for RL training, showing that it can effectively encode verifiable rewards by outperforming a conventional RM in math reasoning. Finally, we provide theoretical guarantees that the variational objective admits a global interior optimum with finite confidence, and we analyze how KL regularization mitigates reward over-optimization.

奖励模型贝叶斯方法强化学习偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。