构建数学工程推理过程-结果对齐评测基准,提升验证器可靠性
PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering
- 设计过程-结果对齐评测基准PRIME,覆盖2530个高难度理工科题目
- 现有验证器常忽略推导错误,新方法使Qwen3模型性能提升超7%
- 验证器在PRIME上的准确率可预测强化学习训练效果,相关性超0.92
基于模型的验证器对于扩展可验证奖励强化学习(RLVR)至关重要,但当前以结果为中心的验证范式主要关注最终答案与真实值的一致性,常忽略推导过程中的潜在错误,导致正确答案但错误推导仍获正向奖励。为弥合这一差距,我们提出PRIME,一个用于评估数学与工程领域推理过程-结果对齐验证能力的基准。该基准从大学级STEM问题中精选,经一致性过滤后共包含2,530个高难度样本。大量实验表明,当前验证器频繁无法检测推导缺陷。我们进一步提出一种基于过程感知的RLVR训练范式,采用通过PRIME筛选的验证器。该方法显著优于仅依赖结果的基线,在AIME24、AIME25和Beyond-AIME上分别实现8.29%、9.12%和7.31%的绝对性能提升,针对Qwen3-14B-Base模型。最后,我们证明了验证器在PRIME上的准确率与RLVR训练效果之间存在强线性相关性(R² > 0.92),验证了PRIME作为验证器选型可靠指标的有效性。
原文摘要 · Abstract (English)
While model-based verifiers are essential for scaling Reinforcement Learning with Verifiable Rewards (RLVR), current outcome-centric verification paradigms primarily focus on the consistency between the final result and the ground truth, often neglecting potential errors in the derivation process. This leads to assigning positive rewards to correct answers produced from incorrect derivations. To bridge this gap, we introduce PRIME, a benchmark for evaluating verifiers on Process-Outcome Alignment verification in Mathematics and Engineering. Curated from a comprehensive collection of college-level STEM problems, PRIME comprises 2,530 high-difficulty samples through a consistency-based filtering pipeline. Through extensive evaluation, we find that current verifiers frequently fail to detect derivation flaws. Furthermore, we propose a process-aware RLVR training paradigm utilizing verifiers selected via PRIME. This approach substantially outperforms the outcome-only verification baseline, achieving absolute performance gains of 8.29%, 9.12%, and 7.31% on AIME24, AIME25, and Beyond-AIME, respectively, for the Qwen3-14B-Base model. Finally, we demonstrate a strong linear correlation ($R^2 > 0.92$) between verifier accuracy on PRIME and RLVR training effectiveness, validating PRIME as a reliable predictor for verifier selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。