arXiv:2603.00895cs.LG2026-03被引 1

用AI批改真实手写微积分作业,准确率接近助教水平。

Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark

  • 基于OCR+大模型,按评分标准生成分数和反馈
  • 90%以上反馈被评正确或可接受,与助教打分高度一致
  • 为手写数学题AI评分提供可复现的基准测试框架

大型本科理工课程中,因教学负担重,作业批改常缺乏有效反馈。本研究在加州大学欧文分校开展大规模实证实验,评估AI对真实手写单变量微积分题目的评分表现。采用结合OCR与大语言模型、受评分标准引导的提示策略,系统处理近800名学生提交的数千份自由作答测验。在无唯一真值标签的情境下,通过助教评分、学生问卷及独立人工评审进行评估,结果显示AI评分与助教评分高度一致,且多数生成反馈被评定为正确或可接受。该研究揭示了OCR条件下数学推理与部分赋分的核心挑战,分析关键失败模式,提出可操作的评分标准与提示设计原则,并建立多视角评估协议,支持真实课程部署。基于所构建的数据集与评估框架,本文进一步提出标准化手写数学题AI评分基准,以促进可复现比较与后续研究。

原文摘要 · Abstract (English)

Grading in large undergraduate STEM courses often yields minimal feedback due to heavy instructional workloads. We present a large-scale empirical study of AI grading on real, handwritten single-variable calculus work from UC Irvine. Using OCR-conditioned large language models with structured, rubric-guided prompting, our system produces scores and formative feedback for thousands of free-response quiz submissions from nearly 800 students. In a setting with no single ground-truth label, we evaluate performance against official teaching-assistant grades, student surveys, and independent human review, finding strong alignment with TA scoring and a large majority of AI-generated feedback rated as correct or acceptable across quizzes. Beyond calculus, this setting highlights core challenges in OCR-conditioned mathematical reasoning and partial-credit assessment. We analyze key failure modes, propose practical rubric- and prompt-design principles, and introduce a multi-perspective evaluation protocol for reliable, real-course deployment. Building on the dataset and evaluation framework developed here, we outline a standardized benchmark for AI grading of handwritten mathematics to support reproducible comparison and future research.

AI批改手写识别教育评测大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。