arXiv:2505.15801cs.CLcs.AI2025-05被引 17

构建验证基准,评估大模型推理训练中的参考答案评分系统。

VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models

  • 设计两个针对参考答案评分的评测基准,含难易变体。
  • 大模型验证器在常规任务表现好,难题仍差得远。
  • 适合研究强化学习推理模型和评分机制的学者。

大型推理模型如OpenAI o1和DeepSeek-R1在复杂推理任务中表现出色,其训练关键在于强化学习中引入基于参考答案的奖励系统,通过对比模型输出与标准答案进行评估。然而,现有奖励评测侧重于响应间的偏好比较,缺乏对实际参考答案验证能力的评估,导致训练中关键环节的评价缺失。本文提出VerifyBench及其挑战版VerifyBench-Hard,两个专为评估参考答案驱动奖励系统而设计的基准。通过精心的数据采集与人工标注,确保高质量。全面评估显示,尽管大模型验证器在标准任务中展现潜力,但在高难度实例上仍存在显著提升空间。通过对不同推理任务和错误类型的表现分析,揭示改进方向。该基准为提升验证准确性提供标准化框架,最终增强强化学习训练模型的推理能力。

原文摘要 · Abstract (English)

Large reasoning models such as OpenAI o1 and DeepSeek-R1 have demonstrated remarkable performance in complex reasoning tasks. A critical component of their training is the incorporation of reference-based reward systems within reinforcement learning (RL), where model outputs are evaluated against ground truth references. However, existing reward benchmarks focus on preference comparisons between responses rather than evaluating verification against ground truth references, leaving a critical gap in our ability to evaluate verification systems used in reasoning model training. In this paper, we introduce VerifyBench and its challenging variant VerifyBench-Hard, two benchmarks specifically designed to assess reference-based reward systems. These benchmarks are constructed through meticulous data collection and curation, followed by careful human annotation to ensure high quality. Our comprehensive evaluation reveals that while larger model-based verifiers show promise on standard cases, all current systems demonstrate substantial room for improvement on challenging instances. Through systematic analysis of performance patterns across reasoning tasks and error categories, we provide insights for advancing reference-based reward systems. These benchmarks establish a standardized framework for improving verification accuracy, ultimately enhancing reasoning capabilities in models trained via RL.

大模型推理奖励系统评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。