评测大模型对数学奥赛证明题的评分能力,提出自动评分新方法。
CombiGraph-Vis: A Curated Multimodal Olympiad Benchmark for Discrete Mathematical Reasoning
- 设计多步代理流程,自动提取参考解法并生成题目专属评分标准。
- 在90份生成解法上,模型对错误的检出率高,但部分分分配存在偏差。
- 适用于需要精细评分的数学推理任务,尤其适合评测和改进AI评分系统。
当前最先进的大语言模型已能解决2025年国际数学奥林匹克(IMO)大部分题目,部分系统声称可完成6道题中的5道。在此背景下,我们评估这些模型对证明题的评分能力:能否准确识别错误、判断严重程度,并合理分配部分分数,而不仅限于正确/错误二元判定。我们构建了一个包含90份Gemini 2.5 Pro生成解法的语料库,采用1-4分制进行人工标注并附详细错误注释;同时使用MathArena提供的IMO/USAMO 2025解法,按0-7分制评分。分析显示,模型能可靠检测出错误(包括细微错误),但在部分分分配上存在校准偏差。为此,我们提出基于代理的工作流,自动提取参考解法并生成针对具体题目的评分细则,实现多步评分。我们对比了不同设计选择,评估其权衡。在所建语料库与MathArena数据集上,所提工作流在与人类评分的一致性及部分分处理的稳定性方面均表现更优。所有代码、数据、提示词与日志均已开源,以促进后续研究。
原文摘要 · Abstract (English)
State-of-the-art (SOTA) LLMs have progressed from struggling on proof-based Olympiad problems to solving most of the IMO 2025 problems, with leading systems reportedly handling 5 of 6 problems. Given this progress, we assess how well these models can grade proofs: detecting errors, judging their severity, and assigning fair scores beyond binary correctness. We study proof-analysis capabilities using a corpus of 90 Gemini 2.5 Pro-generated solutions that we grade on a 1-4 scale with detailed error annotations, and on MathArena solution sets for IMO/USAMO 2025 scored on a 0-7 scale. Our analysis shows that models can reliably flag incorrect (including subtly incorrect) solutions but exhibit calibration gaps in how partial credit is assigned. To address this, we introduce agentic workflows that extract and analyze reference solutions and automatically derive problem-specific rubrics for a multi-step grading process. We instantiate and compare different design choices for the grading workflows, and evaluate their trade-offs. Across our annotated corpus and MathArena, our proposed workflows achieve higher agreement with human grades and more consistent handling of partial credit across metrics. We release all code, data, and prompts/logs to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。