arXiv:2603.02225cs.LG2026-03

无需人工标注,用网页文本训练奖励模型,显著提升数学与安全性能。

Scaling Reward Modeling without Human Supervision

  • 基于网页文本前缀后缀做偏好学习,实现无监督奖励建模。
  • 在1100万数学相关文本上训练,平均提升RewardBench v2准确率7.7点。
  • 适用于不同模型规模与架构,对数学任务和安全评估均有明显增益。

从反馈中学习是推动前沿模型能力与安全性的关键过程,但常受限于成本与可扩展性。本文开展一项试点研究,探索无监督方式下奖励模型的规模化。我们将奖励建模规模化(RBS)最简形式定义为对大规模网络语料中文档前缀与后缀的偏好学习。该方法在多个方面展现优势:尽管未使用任何人工标注,仅在1100万数学相关网页数据上训练,即可在RewardBench v1和v2上持续取得提升,且性能改善可稳定迁移至不同初始化主干模型,涵盖多种模型家族与规模。在各类模型上,该方法使RewardBench v2平均提升7.7点,数学子集最高达+16.1,跨域安全与通用子集亦有持续改进。应用于最佳N选一与策略优化时,该奖励模型显著提升下游数学表现,达到或超过同等规模的监督基线模型。总体而言,我们证明了无需昂贵且不可靠的人工标注即可训练奖励模型的可行性与潜力。

原文摘要 · Abstract (English)

Learning from feedback is an instrumental process for advancing the capabilities and safety of frontier models, yet its effectiveness is often constrained by cost and scalability. We present a pilot study that explores scaling reward models through unsupervised approaches. We operationalize reward-based scaling (RBS), in its simplest form, as preference learning over document prefixes and suffixes drawn from large-scale web corpora. Its advantage is demonstrated in various aspects: despite using no human annotations, training on 11M tokens of math-focused web data yields steady gains on RewardBench v1 and v2, and these improvements consistently transfer across diverse initialization backbones spanning model families and scales. Across models, our method improves RewardBench v2 accuracy by up to +7.7 points on average, with gains of up to +16.1 on in-domain math subsets and consistent improvements on out-of-domain safety and general subsets. When applied to best-of-N selection and policy optimization, these reward models substantially improve downstream math performance and match or exceed strong supervised reward model baselines of similar size. Overall, we demonstrate the feasibility and promise of training reward models without costly and potentially unreliable human annotations.

奖励建模无监督学习数学推理可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。