arXiv:2507.22958cs.CVcs.AI2025-07被引 2

用AI自动批改俄式高考数学手写答案,发现当前模型仍不靠谱。

CHECK-MAT: Checking Hand-Written Mathematical Answers for the Russian Unified State Exam

  • 构建新评测基准,专注判断学生解题思路与错误
  • 测试7个主流大模型在122份真实考卷上的评分准确率
  • 揭示现有模型在数学推理与人工评分标准间存在明显差距

本文提出一个新型评测基准EGE-Math Solutions Assessment Benchmark,用于评估视觉语言模型(VLMs)在判别手写数学解答方面的能力。不同于以往聚焦于解题的评测,本工作关注模型理解学生作答、识别错误并依据固定评分标准打分的能力。研究收集了122份来自俄罗斯统一国家考试(EGE)的扫描答卷及官方专家评分,并在三种推理模式下评估了来自Google、OpenAI、Arcee AI和阿里云的七款现代VLMs。结果表明,当前模型在数学推理能力和与人工评分标准的一致性方面仍存在显著局限,为人工智能辅助评估开辟了新的研究方向。代码已开源:https://github.com/Karifannaa/Auto-check-EGE-math

原文摘要 · Abstract (English)

This paper introduces a novel benchmark, EGE-Math Solutions Assessment Benchmark, for evaluating Vision-Language Models (VLMs) on their ability to assess hand-written mathematical solutions. Unlike existing benchmarks that focus on problem solving, our approach centres on understanding student solutions, identifying mistakes, and assigning grades according to fixed criteria. We compile 122 scanned solutions from the Russian Unified State Exam (EGE) together with official expert grades, and evaluate seven modern VLMs from Google, OpenAI, Arcee AI, and Alibaba Cloud in three inference modes. The results reveal current limitations in mathematical reasoning and human-rubric alignment, opening new research avenues in AI-assisted assessment. You can find code in https://github.com/Karifannaa/Auto-check-EGE-math

AI批改数学评估视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。