用AI批改真实手写微积分作业,准确率接近助教水平。
Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark
- 基于OCR+大模型,按评分标准生成分数和反馈
- 90%以上反馈被评正确或可接受,与助教打分高度一致
- 为手写数学题AI评分提供可复现的基准测试框架
大型本科理工课程中,因教学负担重,作业批改常缺乏有效反馈。本研究在加州大学欧文分校开展大规模实证实验,评估AI对真实手写单变量微积分题目的评分表现。采用结合OCR与大语言模型、受评分标准引导的提示策略,系统处理近800名学生提交的数千份自由作答测验。在无唯一真值标签的情境下,通过助教评分、学生问卷及独立人工评审进行评估,结果显示AI评分与助教评分高度一致,且多数生成反馈被评定为正确或可接受。该研究揭示了OCR条件下数学推理与部分赋分的核心挑战,分析关键失败模式,提出可操作的评分标准与提示设计原则,并建立多视角评估协议,支持真实课程部署。基于所构建的数据集与评估框架,本文进一步提出标准化手写数学题AI评分基准,以促进可复现比较与后续研究。
原文摘要 · Abstract (English)
Grading in large undergraduate STEM courses often yields minimal feedback due to heavy instructional workloads. We present a large-scale empirical study of AI grading on real, handwritten single-variable calculus work from UC Irvine. Using OCR-conditioned large language models with structured, rubric-guided prompting, our system produces scores and formative feedback for thousands of free-response quiz submissions from nearly 800 students. In a setting with no single ground-truth label, we evaluate performance against official teaching-assistant grades, student surveys, and independent human review, finding strong alignment with TA scoring and a large majority of AI-generated feedback rated as correct or acceptable across quizzes. Beyond calculus, this setting highlights core challenges in OCR-conditioned mathematical reasoning and partial-credit assessment. We analyze key failure modes, propose practical rubric- and prompt-design principles, and introduce a multi-perspective evaluation protocol for reliable, real-course deployment. Building on the dataset and evaluation framework developed here, we outline a standardized benchmark for AI grading of handwritten mathematics to support reproducible comparison and future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。