arXiv:2606.11477cs.CVcs.AI2026-06

用视觉语言模型实现公平高效的纸质答题自动评分

Towards Fully Automated Exam Grading: Fairness-Aware Recognition of Handwritten Answers with Foundation Models

论文配图:Towards Fully Automated Exam Grading: Fairness-Aware Recognition of Handwritten Answers with Foundation Models
图 1 · 摘自论文原文
  • 利用通用视觉语言模型理解手写答案,而非简单模板匹配
  • 在61份匿名试卷上达到98.4%准确率,误判率降至0.58%
  • 重点优化对错判的公平性,适合大规模教育评估场景

人工批改手写试卷耗时且易出错,尤其面对大规模学生群体;而完全数字化考试则限制了开放题形式。一种折中方案是保留纸质答题,仅将关键答案以大写字母形式填入表格,由机器读取。当前自动识别准确率仅为88%–91%,且难以处理答案超出单元格、划掉或草书等情况。本文证明,通用视觉-语言基础模型(VLMs)能有效理解页面语义,显著提升识别性能。在包含61份匿名试卷(共3141个作答位置)的基准测试中,最佳模型达到98.4%准确率。关键在于公平性评估:区分对学生不利的漏判(正确答案被判错误)与误判,并通过轻量提示引入标准答案作为上下文,使漏判率降至0.58%。在示例评分方案下,仅有3份试卷被误判,且均可通过学生自审发现。因此,大规模、公平感知的全自动阅卷已具备可行性,研究团队公开了匿名数据集以支持可复现性。

原文摘要 · Abstract (English)

Correcting handwritten exams by hand is time-consuming and error-prone, particularly for large cohorts, while fully digital exams tend to force a didactic narrowing towards closed question formats. A practical middle ground keeps paper-based, problem-oriented tasks but records the assessment-relevant answers as single capital letters in a table that a machine can read. The open question is whether this reading can be made accurate and, above all, fair enough for unsupervised grading. Earlier automated approaches reached only about 88%--91% recognition -- too low -- and failed on the cases that matter most: answers placed outside the cell, crossed out, or written in cursive. We show that general-purpose vision-language foundation models (VLMs), which interpret the page rather than match pixel templates, close this gap. On a benchmark of 61 anonymised exams (3141 answer positions) the best model reaches 98.4% accuracy, well above the previous baseline. Crucially, we centre the evaluation on fairness: we distinguish false negatives (a correct answer marked wrong, which disadvantages the student) from false positives, and a lightweight prompt that supplies the reference solution as context lowers the false-negative rate to 0.58%. Under an exemplary grading scheme only three of the 61 exams would be graded worse, all caught by a student self-review step. Fully automated, fairness-aware exam grading at scale is therefore defensible; we release the anonymised benchmark to support reproducibility.

自动阅卷视觉语言模型教育公平手写识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。