arXiv:2602.06291cs.CL2026-02被引 1

不用专家验证,通过解题后果评估数学推理质量。

Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math

  • 用相关题目验证候选解的实用价值,替代人工判断。
  • 在10个问题上,正确率提升9.1点,AUC提升8.2点。
  • 适合评估大模型数学推理能力,尤其对失败案例仍有效。

近期推理模型进展表明,生成研究级数学问题的合理解法可能已可实现,但验证仍是瓶颈,耗时且依赖专家。我们提出一种基于后果的评估方法——后果效用(Consequence-Based Utility),其核心思想是:一个有意义的解法应包含足够的方法层面信息,使其在解决邻近相关问题时表现优于错误解法。该方法无需人工标注,通过测试候选解作为上下文示例在相关可验证问题上的求解效果来打分。我们在一组原创的研究级数学问题上进行了评估,每个问题配有一份专家解法和九份大模型生成解法。结果表明,该方法在排序质量上持续优于奖励模型、生成式奖励模型及大模型裁判。例如,在GPT-OSS-120B上,准确率从67.2%提升至76.3%,AUC从71.4%提升至79.6%;在GPT-OSS-20B上,AUC也从69.0%提升至79.2%。此外,相比大模型裁判,该方法表现出更大的求解器-评估者差距,即使在求解器常失败的样本上,仍能保持较好的正确与错误区分能力。

原文摘要 · Abstract (English)

Recent progress in reasoning models suggests that generating plausible attempts for research-level mathematics may be within reach, but verification remains a bottleneck, consuming scarce expert time. We hypothesize that a meaningful solution should contain enough method-level information that, when applied to a neighborhood of related questions, it should yield better downstream performance than incorrect solutions. Building on this idea, we propose \textbf{Consequence-Based Utility}, an oracle-free evaluator that scores each candidate by testing its value as an in-context exemplar in solving related yet verifiable questions. Our approach is evaluated on an original set of research-level math problems, each paired with one expert-written solution and nine LLM-generated solutions. Notably, Consequence-Based Utility consistently outperforms reward models, generative reward models, and LLM judges on ranking quality. Specifically, for GPT-OSS-120B, it improves Acc@1 from 67.2 to 76.3 and AUC from 71.4 to 79.6, with similarly large AUC gains on GPT-OSS-20B (69.0 to 79.2). Furthermore, compared to LLM-Judges, it also exhibits a larger solver-evaluator gap, maintaining a stronger correct-wrong separation even on instances where the underlying solver often fails to solve.

数学推理模型评估无监督评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。