arXiv:2601.20327cs.CL2026-01ACL

用两阶段滚动和统一标准训练生成式奖励模型,提升文本生成的强化学习效果。

CE-RM: A Pointwise Generative Reward Model Optimized via Two-Stage Rollout and Unified Criteria

  • 基于两阶段滚动与统一查询标准训练点对点生成奖励模型
  • 仅用5.7K高质量数据即在多种评测中表现更优,尤其在Best-of-N场景
  • 适合需要高精度自动评估的生成式AI研究者与工程师

自动评估对开放域自然语言生成至关重要但极具挑战性,尤其当基于规则的度量不可行时。相比传统方法,近期的大型语言模型作为裁判范式展现出更优且灵活的评估能力,有潜力作为强化学习中的生成式奖励模型。然而,已有研究发现其在基准测试中的优异表现与其在强化学习实践中的实际有效性之间存在显著差距。我们将其归因于现有研究的若干局限,包括成对评估占主导地位以及评估标准优化不足。为此,我们提出CE-RM-4B,一个通过专用两阶段滚动方法训练、采用统一查询标准的点对点生成式奖励模型。仅使用约5.7K来自开源偏好数据集的高质量数据,我们的模型在多种奖励模型基准测试中取得更优表现,特别是在Best-of-N场景下,并在下游强化学习实践中带来更有效的性能提升。

原文摘要 · Abstract (English)

Automatic evaluation is crucial yet challenging for open-ended natural language generation, especially when rule-based metrics are infeasible. Compared with traditional methods, the recent LLM-as-a-Judge paradigms enable better and more flexible evaluation, and show promise as generative reward models for reinforcement learning. However, prior work has revealed a notable gap between their seemingly impressive benchmark performance and actual effectiveness in RL practice. We attribute this issue to some limitations in existing studies, including the dominance of pairwise evaluation and inadequate optimization of evaluation criteria. Therefore, we propose CE-RM-4B, a pointwise generative reward model trained with a dedicated two-stage rollout method, and adopting unified query-based criteria. Using only about 5.7K high-quality data curated from the open-source preference dataset, our CE-RM-4B achieves superior performance on diverse reward model benchmarks, especially in Best-of-N scenarios, and delivers more effective improvements in downstream RL practice.

生成式奖励模型强化学习自动评估文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。