构建跨领域验证基准,系统评估各类推理验证器性能优劣。
VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
- 设计四维实验框架,对比专用验证器与通用大模型在不同输入条件下的表现。
- 4000道专家级题目显示:专用验证器准确率高但召回低,通用模型包容性强但精度不稳。
- 揭示验证器对输入结构敏感及跨领域泛化能力差的核心瓶颈,指导强化学习奖励设计。
大型语言模型(LLMs)越来越多地依赖强化学习(RL)通过反馈提升推理能力。一个关键挑战是验证模型生成答案与参考答案的一致性,因为这些答案往往冗长、多样且微妙。基于规则的验证器难以应对复杂情况,促使采用基于模型的验证器。然而,专用验证器灵活性不足,而通用大模型评判者又存在不一致性。现有研究主要致力于构建更好的验证器,但缺乏对不同类型验证器在多个领域中性能的系统评估,严重制约了可验证奖励强化学习(RLVR)的可靠发展。为此,我们提出VerifyBench——一个跨领域的综合性基准,用于系统评估验证器性能。我们构建了4000道涵盖数学、物理、化学和生物的专家级问题,每个问题配有参考答案和多样化生成回答。评估可靠性通过多学科专家团队的严格标注流程保障。我们设计了一个四维实验框架,全面比较专用验证器与通用大模型在提取答案与完整回答、短输出与长输出组合条件下的性能边界。评估揭示了验证器的根本权衡:专用验证器虽达领先准确率,却存在召回缺陷;通用模型展现更强包容性但精度不稳定。更重要的是,我们发现验证器对输入结构高度敏感,且在跨领域泛化方面存在固有局限,为当前验证器技术瓶颈提供了关键洞见。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly rely on reinforcement learning (RL) to enhance their reasoning capabilities through feedback. A critical challenge is verifying the consistency of model-generated responses and reference answers, since these responses are often lengthy, diverse, and nuanced. Rule-based verifiers struggle with complexity, prompting the use of model-based verifiers. However, specialized verifiers lack flexibility, while general LLM judges can be inconsistent. Existing research primarily focuses on building better verifiers, yet a systematic evaluation of different types of verifiers' performance across domains remains lacking, severely constraining the reliable development of Reinforcement Learning with Verifiable Reward (RLVR). To address this, we propose VerifyBench--a cross-domain comprehensive benchmark for systematically evaluating verifiers. We construct 4,000 expert-level questions covering mathematics, physics, chemistry, and biology. Each question is equipped with reference answers and diverse responses. The reliability of the evaluation is ensured through a rigorous annotation process conducted by a multidisciplinary expert team. We design a four-dimensional experimental framework to comprehensively compare the performance boundaries of specialized verifiers and general LLMs under combined conditions of extracted answers vs. complete responses, and short vs. long outputs. Our evaluation uncovers fundamental trade-offs in verifiers: while specialized verifiers achieve leading accuracy, they exhibit deficiencies in recall; general models show stronger inclusivity but unstable precision. More importantly, we discover verifiers' high sensitivity to input structure and inherent limitations in cross-domain generalization, providing critical insights into the bottlenecks of current verifier technology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。