构建数学推理验证基准,评估模型每一步的逻辑正确性。
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
- 人工标注超500小时,构建细粒度步骤级验证数据集。
- 29个生成式评判模型中多数开源模型表现远逊闭源模型。
- 揭示验证能力瓶颈,探索自验证与生成动态机制。
基于大语言模型的推理系统近期在IMO 2025竞赛中达到金牌水平,其数学证明需每一步均正确且充分支持方可获得满分。为训练此类开放性前沿数学任务中的推理模型,具备捕捉步骤级错误的强验证器至关重要。我们提出Hard2Verify,一个由超过500小时人工标注构建的步骤级验证基准。该基准旨在严格评估前沿验证器:要求验证器对前沿大模型生成的复杂、开放性数学问题回答进行步骤级标注或识别首个错误。我们评估了29个生成式批评者及处理奖励模型,发现除少数优秀模型外,开源验证器普遍落后于闭源模型。随后分析了步骤级验证性能差的原因、验证器算力扩展的影响,以及自验证与验证-生成动态等根本问题。
原文摘要 · Abstract (English)
Large language model (LLM)-based reasoning systems have recently achieved gold medal-level performance in the IMO 2025 competition, writing mathematical proofs where, to receive full credit, each step must be not only correct but also sufficiently supported. To train LLM-based reasoners in such challenging, open-ended settings, strong verifiers capable of catching step-level mistakes are necessary prerequisites. We introduce Hard2Verify, a human-annotated, step-level verification benchmark produced with over 500 hours of human labor. Hard2Verify is designed to rigorously assess step-level verifiers at the frontier: Verifiers must provide step-level annotations or identify the first error in responses generated by frontier LLMs for very recent, challenging, and open-ended math questions. We evaluate 29 generative critics and process reward models, demonstrating that, beyond a few standouts, open-source verifiers lag closed source models. We subsequently analyze what drives poor performance in step-level verification, the impacts of scaling verifier compute, as well as fundamental questions such as self-verification and verification-generation dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。