arXiv:2603.13083cs.CYcs.AI2026-03被引 2

用人工介入的LLM系统高效批改手写数学作业,兼顾速度与准确。

Human-in-the-Loop LLM Grading for Handwritten Mathematics Assessments

  • 构建解题标准与评分细则,指导LLM分步打分
  • 减少23%阅卷时间,评分一致性优于或等同人工
  • 混合设计确保错误可控,适合小测验场景

及时且个性化的手写作业反馈对学习大有裨益,但大规模实施困难。随着生成式AI削弱了课后测评的可靠性,教学更倾向于监督式课堂评估。本文提出一种可扩展的端到端工作流,用于辅助低分值课堂测试的手写数学作业评分。流程包括:(1) 构建标准答案;(2) 制定基于评分表的详细评分规则以引导LLM;(3) 结合自动扫描与匿名化、多轮LLM评分、自动一致性检测及强制人工复核的评分流程。我们在两门本科数学课程中部署该系统,使用六次低风险课堂测试进行验证。实证结果表明,使用LLM辅助使阅卷时间减少约23%,评分一致性与全人工评分相当,部分情况甚至更优。模型偶发错误,但通过混合设计有效控制。总体显示,精心设计的人机协同评分能显著减轻负担,同时保障公平与准确。

原文摘要 · Abstract (English)

Providing timely and individualised feedback on handwritten student work is highly beneficial for learning but difficult to achieve at scale. This challenge has become more pressing as generative AI undermines the reliability of take-home assessments, shifting emphasis toward supervised, in-class evaluation. We present a scalable, end-to-end workflow for LLM-assisted grading of short, pen-and-paper assessments. The workflow spans (1) constructing solution keys, (2) developing detailed rubric-style grading keys used to guide the LLM, and (3) a grading procedure that combines automated scanning and anonymisation, multi-pass LLM scoring, automated consistency checks, and mandatory human verification. We deploy the system in two undergraduate mathematics courses using six low-stakes in-class tests. Empirically, LLM assistance reduces grading time by approximately 23% while achieving agreement comparable to, and in several cases tighter than, fully manual grading. Occasional model errors occur but are effectively contained by the hybrid design. Overall, our results show that carefully embedded human-in-the-loop LLM grading can substantially reduce workload while maintaining fairness and accuracy.

数学教育LLM评分人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。