arXiv:2603.12963cs.CL2026-03AAAI被引 2

首个专为长文本生成设计的奖励模型评测基准。

Long-form RewardBench: Evaluating Reward Models for Long-form Generation

  • 构建五类长文本任务,系统评估奖励模型能力。
  • 发现现有模型在长文本奖励上表现不足,且错误位置影响评分。
  • 分类器比生成式模型更具泛化性,适合长文本场景。

强化学习对齐的广泛应用凸显了奖励模型的重要性。尽管已有多个领域和场景的评测基准,但针对长文本生成的奖励模型评估仍存在显著空白。为此,我们提出Long-form RewardBench,首个专为长文本生成设计的奖励模型评测平台。该基准包含问答、RAG、对话、写作和推理五类核心子任务。通过多阶段精心设计的数据收集流程,获取指令与偏好数据,并在20余种主流奖励模型(包括分类器与生成式模型)上开展广泛实验。结果表明,当前模型仍缺乏长文本奖励建模能力。我们设计了新颖的长文本“针在草堆中”测试,揭示奖励模型表现与错误位置及响应长度相关,且分类器与生成式模型表现出不同特征。最后,我们证明分类器在相同数据下具备更强泛化能力。作为首个长文本奖励建模基准,本工作旨在为该关键领域提供稳健的进展可视化平台。

原文摘要 · Abstract (English)

The widespread adoption of reinforcement learning-based alignment highlights the growing importance of reward models. Various benchmarks have been built to evaluate reward models in various domains and scenarios. However, a significant gap remains in assessing reward models for long-form generation, despite its critical role in real-world applications. To bridge this, we introduce Long-form RewardBench, the first reward modeling testbed specifically designed for long-form generation. Our benchmark encompasses five key subtasks: QA, RAG, Chat, Writing, and Reasoning. We collected instruction and preference data through a meticulously designed multi-stage data collection process, and conducted extensive experiments on 20+ mainstream reward models, including both classifiers and generative models. Our findings reveal that current models still lack long-form reward modeling capabilities. Furthermore, we designed a novel Long-form Needle-in-a-Haystack Test, which revealed a correlation between reward modeling performance and the error's position within a response, as well as the overall response length, with distinct characteristics observed between classification and generative models. Finally, we demonstrate that classifiers exhibit better generalizability compared to generative models trained on the same data. As the first benchmark for long-form reward modeling, this work aims to offer a robust platform for visualizing progress in this crucial area.

奖励模型长文本生成评测基准RLHF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。