用智能体流程自动批改数学竞赛证明,提升评分准确性与一致性
RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows
- 设计多步智能体流程,解析参考答案并生成题目专属评分标准
- 在IMO/USAMO 2025数据集上,模型与人工评分一致率达87%以上
- 适合需要高精度自动化批改的数学竞赛评测与教育AI研究者
当前顶级大模型已能解决多数国际数学奥林匹克(IMO)2025题,部分系统可完成6题中的5题。在此背景下,我们评估大模型在证明题评分上的能力:识别错误、判断严重程度,并给出非二元正确性分数。我们构建了90份由Gemini 2.5 Pro生成的解题方案,采用1-4分制标注并附详细错误信息;同时使用MathArena提供的IMO/USAMO 2025解题集,按0-7分制评分。分析表明,模型可可靠识别错误(包括细微错误),但在部分分分配上存在校准偏差。为此,我们提出基于智能体的工作流,通过提取参考解法并自动生成题目标杆评分细则,实现多阶段评分。我们在多个设计选项中进行比较与评估,结果显示所提方法在标注语料和MathArena上均显著提升与人工评分的一致性,并更稳定地处理部分分。所有代码、数据、提示词及日志均已公开。
原文摘要 · Abstract (English)
State-of-the-art (SOTA) LLMs have progressed from struggling on proof-based Olympiad problems to solving most of the IMO 2025 problems, with leading systems reportedly handling 5 of 6 problems. Given this progress, we assess how well these models can grade proofs: detecting errors, judging their severity, and assigning fair scores beyond binary correctness. We study proof-analysis capabilities using a corpus of 90 Gemini 2.5 Pro-generated solutions that we grade on a 1-4 scale with detailed error annotations, and on MathArena solution sets for IMO/USAMO 2025 scored on a 0-7 scale. Our analysis shows that models can reliably flag incorrect (including subtly incorrect) solutions but exhibit calibration gaps in how partial credit is assigned. To address this, we introduce agentic workflows that extract and analyze reference solutions and automatically derive problem-specific rubrics for a multi-step grading process. We instantiate and compare different design choices for the grading workflows, and evaluate their trade-offs. Across our annotated corpus and MathArena, our proposed workflows achieve higher agreement with human grades and more consistent handling of partial credit across metrics. We release all code, data, and prompts/logs to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。