arXiv:2510.05538cs.CVcs.AI2025-10被引 1

测试多模态大模型批改手写数学作业能力,发现对算式有效但对画图题表现差。

Seeing the Big Picture: Evaluating Multimodal LLMs' Ability to Interpret and Grade Handwritten Student Work

  • 用真实手写作业测试模型批改能力,分算术题和画图题两组实验。
  • 算术题准确率达95%,但画图题仅0.20的评分一致性,提升描述后升至0.47。
  • 适合教育科技研究者与一线教师关注模型在真实教学场景的局限性。

多模态大语言模型在批改手写学生作业方面展现出潜力,尤其在小学和初中数学教育中,手写过程能反映学习思维,但人工批改耗时。我们开展两项实验:实验A评估加纳初中生288份算术题手写答案,模型达到95%准确率(k=0.90),接近人类水平;实验B分析美国小学生150份数学图画作答,题目无唯一答案,需结合视觉理解与教学判断。模型直接分析图像时,评分一致性仅k=0.20,而提供详细人工描述后,一致性提升至k=0.47,接近人与人之间的一致性。结果表明,模型能较好理解算术表达,但在解读学生数学图示方面仍有显著不足。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) raise the question of their potential for grading, analyzing, and offering feedback on handwritten student classwork. This capability would be particularly beneficial in elementary and middle-school mathematics education, where most work remains handwritten, because seeing students' full working of a problem provides valuable insights into their learning processes, but is extremely time-consuming to grade. We present two experiments investigating MLLM performance on handwritten student mathematics classwork. Experiment A examines 288 handwritten responses from Ghanaian middle school students solving arithmetic problems with objective answers. In this context, models achieved near-human accuracy (95%, k = 0.90) but exhibited occasional errors that human educators would be unlikely to make. Experiment B evaluates 150 mathematical illustrations from American elementary students, where the drawings are the answer to the question. These tasks lack single objective answers and require sophisticated visual interpretation as well as pedagogical judgment in order to analyze and evaluate them. We attempted to separate MLLMs' visual capabilities from their pedagogical abilities by first asking them to grade the student illustrations directly, and then by augmenting the image with a detailed human description of the illustration. We found that when the models had to analyze the student illustrations directly, they struggled, achieving only k = 0.20 with ground truth scores, but when given human descriptions, their agreement levels improved dramatically to k = 0.47, which was in line with human-to-human agreement levels. This gap suggests MLLMs can "see" and interpret arithmetic work relatively well, but still struggle to "see" student mathematical illustrations.

多模态教育AI手写识别模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。