arXiv:2510.24636cs.CL2025-10被引 3

用工具辅助评估长文本任务,让奖励模型更准判别质量差异。

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

  • 引入工具调用获取外部证据,提升对长文本的判断能力。
  • 在27000+样本上训练,显著优于现有奖励模型。
  • 适合需要可靠长文本评估的LLM对齐任务。

奖励模型(RMs)已成为对齐大语言模型的关键,作为训练和推理中人类评估的可扩展替代。然而,现有模型在知识密集型和长文本任务中表现受限,因评估正确性需依赖模型内部知识之外的依据。这一缺陷导致其难以可靠区分细微质量差异,尤其当需外部证据时。为此,我们提出OpenRM,一种通过调用外部工具系统性评估开放性回答的工具增强型长文本奖励模型。我们在超过27,000个合成的成对样本上,使用可控数据合成框架与组相对策略优化(GRPO)训练OpenRM,训练目标同时监督中间工具使用和最终结果准确性,激励模型学习基于证据的判断策略。在三个新收集的数据集及两个常用基准上的大量实验表明,OpenRM显著优于现有奖励建模方法。进一步地,我们将OpenRM应用于推理时响应选择和训练时数据选择,均带来下游大模型对齐任务的持续提升,凸显了工具增强型奖励模型在规模化可靠长文本评估中的潜力。

原文摘要 · Abstract (English)

Reward models (RMs) have become essential for aligning large language models (LLMs), serving as scalable proxies for human evaluation in both training and inference. However, existing RMs struggle on knowledge-intensive and long-form tasks, where evaluating correctness requires grounding beyond the model's internal knowledge. This limitation hinders them from reliably discriminating subtle quality differences, especially when external evidence is necessary. To address this, we introduce OpenRM, a tool-augmented long-form reward model that systematically judges open-ended responses by invoking external tools to gather relevant evidence. We train OpenRM with Group Relative Policy Optimization (GRPO) on over 27K synthesized pairwise examples generated through a controllable data synthesis framework. The training objective jointly supervises intermediate tool usage and final outcome accuracy, incentivizing our reward model to learn effective evidence-based judgment strategies. Extensive experiments on three newly-collected datasets and two widely-used benchmarks demonstrate that OpenRM substantially outperforms existing reward modeling approaches. As a further step, we integrate OpenRM into both inference-time response selection and training-time data selection. This yields consistent gains in downstream LLM alignment tasks, highlighting the potential of tool-augmented reward models for scaling reliable long-form evaluation.

奖励模型长文本评估工具增强LLM对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。