arXiv:2509.02492cs.CLcs.LG2025-09AAAI被引 5

用自训练让奖励模型会推理,生成带理由的评分结果。

GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning

  • 用无标注数据自训练,让模型学会生成奖励理由
  • 在多个任务上优于现有判别和生成模型
  • 可做基础模型,少调参就能用于排序和强化学习

近年来奖励建模进展得益于从任务特定设计转向通用奖励模型。然而,有效构建奖励模型仍面临根本挑战:高度依赖大规模带标签偏好数据。通过在大量无标签数据上预训练提供了一条有前景的路径,但现有方法未能在模型中引入明确的推理能力。为此,我们提出一种自训练方法,利用无标签数据激发奖励模型的奖励推理能力。基于此,我们开发了GRAM-R²,一种生成式奖励模型,不仅能输出偏好标签,还能生成相应的奖励理由。GRAM-R²可作为奖励推理的基础模型,无需或仅需少量微调即可应用于多种任务。在响应排序、任务适应和基于人类反馈的强化学习实验中,GRAM-R²持续表现出色,优于多个强基准模型。

原文摘要 · Abstract (English)

Significant progress in reward modeling over recent years has been driven by a paradigm shift from task-specific designs towards generalist reward models. Despite this trend, developing effective reward models remains a fundamental challenge: the heavy reliance on large-scale labeled preference data. Pre-training on abundant unlabeled data offers a promising direction, but existing approaches fall short of instilling explicit reasoning into reward models. To bridge this gap, we propose a self-training approach that leverages unlabeled data to elicit reward reasoning in reward models. Based on this approach, we develop GRAM-R$^2$, a generative reward model trained to produce not only preference labels but also accompanying reward rationales. GRAM-R$^2$ can serve as a foundation model for reward reasoning and can be applied to a wide range of tasks with minimal or no additional fine-tuning. It can support downstream applications such as response ranking and task-specific reward tuning. Experiments on response ranking, task adaptation, and reinforcement learning from human feedback demonstrate that GRAM-R$^2$ consistently delivers strong performance, outperforming several strong discriminative and generative baselines.

奖励模型生成式自训练推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。