arXiv:2601.07349cs.CL2026-01中稿 · ICML被引 6

用自然语言反馈替代二元判断,提升生成模型奖励建模的准确性

Reward Modeling from Natural Language Human Feedback

  • 用人类批评与模型生成的相似度作为奖励信号
  • 在多个基准上优于仅依赖结果反馈的模型
  • 适合需要高质量过程评估的AI对齐研究

基于可验证奖励的强化学习(RLVR)已成为训练生成式奖励模型(GRM)的主流方法。传统成对比较任务中,GRM生成带有评述和偏好标签的推理链,而RLVR依赖偏好标签作为训练奖励。然而本文表明,这种二分类任务使GRM容易通过猜测正确结果而非合理评述获得成功,导致奖励信号噪声大,影响强化学习效果。为此,我们提出从自然语言人类反馈中进行奖励建模(RM-NLHF),利用自然语言反馈获取过程奖励信号,缓解二元任务中解空间受限的问题。具体地,通过计算模型生成评述与人类评述的相似度作为训练奖励,提供比仅基于结果的监督更准确的信号。此外,针对人类评述难以规模化的问题,我们引入元奖励模型(MetaRM),学习从含人类评述的数据集预测过程奖励,并泛化到无评述数据。多基准实验表明,本方法持续优于仅使用结果奖励训练的先进GRM,验证了自然语言反馈在监督中的优越性。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable reward (RLVR) on preference data has become the mainstream approach for training Generative Reward Models (GRMs). Typically in pairwise rewarding tasks, GRMs generate reasoning chains ending with critiques and preference labels, and RLVR then relies on the correctness of the preference labels as the training reward. However, in this paper, we demonstrate that such binary classification tasks make GRMs susceptible to guessing correct outcomes without sound critiques. Consequently, these spurious successes introduce substantial noise into the reward signal, thereby impairing the effectiveness of reinforcement learning. To address this issue, we propose Reward Modeling from Natural Language Human Feedback (RM-NLHF), which leverages natural language feedback to obtain process reward signals, thereby mitigating the problem of limited solution space inherent in binary tasks. Specifically, we compute the similarity between GRM-generated and human critiques as the training reward, which provides more accurate reward signals than outcome-only supervision. Additionally, considering that human critiques are difficult to scale up, we introduce Meta Reward Model (MetaRM) which learns to predict process reward from datasets with human critiques and then generalizes to data without human critiques. Experiments on multiple benchmarks demonstrate that our method consistently outperforms state-of-the-art GRMs trained with outcome-only reward, confirming the superiority of integrating natural language over binary human feedback as supervision.

奖励建模自然语言反馈强化学习生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。