arXiv:2607.29185cs.CL2026-07

让奖励模型学会生成可优化的推理轨迹,提升评分准确性。

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End

  • 将推理过程作为离散潜变量,端到端学习以匹配最终评分
  • 在分布内和分布外任务上均超越传统与混合型奖励模型
  • 适合需要高精度奖励评估的复杂任务场景

奖励模型(RMs)是通过强化学习对齐大语言模型与人类偏好的核心组件。传统标量奖励模型虽高效且具概率解释性,但依赖表层线索,难以泛化至复杂或分布外(OOD)任务。相反,生成式奖励模型利用详尽推理提升鲁棒性,但其基于自然语言的评分缺乏数值灵活性和概率可解释性。尽管近期方法通过离线多任务学习结合二者,但并无法保证生成的推理轨迹主动服务于下游标量奖励预测。为此,我们提出LatentRM,一种将中间推理轨迹作为离散潜变量学习的框架,显式最大化下游标量奖励的似然性。通过端到端的在线策略优化潜空间,LatentRM紧密耦合深度推理评估与精确打分。在分布内与分布外数据集及RLHF上的广泛验证表明,该方法在从开放对话到复杂推理的任务中,均优于标量、生成式及混合型奖励模型,在偏好建模与策略对齐方面表现更优。

原文摘要 · Abstract (English)

Reward models (RMs) are central to aligning large language models with human preferences via reinforcement learning. Although traditional scalar RMs enable efficient and probabilistic reward modeling, they rely on superficial cues that fail to generalize to complex or out-of-distribution (OOD) tasks. Conversely, generative RMs leverage extensive reasoning to improve robustness on challenging tasks, but their natural language-based scores lack the numerical flexibility and probabilistic interpretability that scalar RMs offer. While recent approaches combine both paradigms through off-policy multi-task learning, such parallel optimization does not guarantee that generated reasoning traces actively align with or benefit downstream scalar reward prediction. To address this mismatch, we propose LatentRM, a reward modeling framework that learns intermediate reasoning traces as discrete latent variables to explicitly maximize the likelihood of downstream scalar rewards. Through on-policy optimization of the latent reasoning space end-to-end, LatentRM tightly couples deep reasoning-based evaluation with precise scoring. Extensive validations on in-distribution and OOD datasets and RLHF show that LatentRM outperforms scalar, generative, and hybrid RMs on preference modeling and policy alignment across tasks ranging from open-ended conversation to complex reasoning.

奖励模型推理轨迹强化学习大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。