用视觉大模型自动批改手写数学题,准确率高但易因识别错误出错
Automated Grading of Handwritten Mathematics Using Vision-Capable LLMs

- 单次调用大模型完成手写内容识别与评分
- 整体准确率高,87%的错误源于识别失败
- 适合教育科技开发者和教师参考
自动化评分系统已广泛应用于多种题型,但手写数学因其多步骤解题过程仍具挑战。具备视觉能力的大语言模型(LLM)为此带来新可能,但其在真实教学场景中的可靠性尚不明确。本文对基于LLM的手写数学评分系统进行实证评估,采用教师定义的评分标准,扩展先前针对文本输入的流程,在单次LLM调用中集成图像转录与评分。评估基于两所大学理工课程的学生作业。对比AI评分与人工标注的基准结果,发现在评分项层面整体准确率较高,其中87%的错误归因于转录失败而非评分规则误用。我们归纳了常见错误类型,包括图像质量差、内容幻觉及等价表达处理错误。研究揭示了基于LLM的批改系统在手写数学上的潜力与局限,为系统设计、提示优化和教育部署提供指导。
原文摘要 · Abstract (English)
Automated grading systems have enabled scalable assessment for many response types, but handwritten mathematics remains a barrier due to the complexity of multi-step solutions. Vision-capable large language models (LLMs) offer new opportunities here, yet their reliability in authentic instructional settings remains poorly understood. We present an empirical evaluation of an LLM-based grader for handwritten mathematical work using instructor-defined rubrics. Extending a prior pipeline for typed responses, we integrate transcription and rubric-based evaluation of photographic submissions within a single LLM call, evaluating on student work from two university STEM courses. Comparing AI grading decisions against human-assigned ground truth at the rubric-item level, we observe high overall accuracy, with most errors -- 87\% in the best model -- attributable to transcription failures rather than rubric misapplication. We categorize common error modes, including image quality issues, hallucinated content, and incorrect handling of equivalent expressions. These findings highlight both the promise and limitations of LLM-based grading for handwritten mathematics, providing guidance for system design, prompt refinement, and deployment in educational settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。