arXiv:2601.00730cs.CV2026-01被引 1

用多模态大模型自动批改手写工程试卷,保留原始书写形式。

Grading Handwritten Engineering Exams with Multimodal Large Language Models

  • 仅需教师手写标准答案和评分规则,自动生成评分流程。
  • 平均误差约8分(满分40),误判率约17%,可审计可解析。
  • 适合需要高效批改手写答题的高校工科课程教师使用。

手写理工科试卷能体现开放性推理和绘图能力,但人工批改耗时且难扩展。本文提出一种端到端工作流,利用多模态大语言模型对扫描后的手写工程测验进行自动评分,保持A4纸张格式与自由手写风格。教师仅需提供手写参考答案(100%)和简短评分规则;参考答案被转化为纯文本摘要以指导评分,不暴露原始图像。通过多阶段设计实现可靠性:格式/存在性检查防止漏评,独立评分器集成,监督者聚合,以及确定性验证模板生成可审计、机器可读报告。在斯洛文尼亚真实课程测验上进行封闭测试,包含手绘电路图。采用GPT-5.2与Gemini-3 Pro作为后端,全管道在满分40分下平均绝对误差约8分,偏差低,手动复核触发率约17%。消融实验表明,简单提示或移除参考答案会显著降低准确率并引入系统性高估,证实结构化提示与参考基准的重要性。

原文摘要 · Abstract (English)

Handwritten STEM exams capture open-ended reasoning and diagrams, but manual grading is slow and difficult to scale. We present an end-to-end workflow for grading scanned handwritten engineering quizzes with multimodal large language models (LLMs) that preserves the standard exam process (A4 paper, unconstrained student handwriting). The lecturer provides only a handwritten reference solution (100%) and a short set of grading rules; the reference is converted into a text-only summary that conditions grading without exposing the reference scan. Reliability is achieved through a multi-stage design with a format/presence check to prevent grading blank answers, an ensemble of independent graders, supervisor aggregation, and rigid templates with deterministic validation to produce auditable, machine-parseable reports. We evaluate the frozen pipeline in a clean-room protocol on a held-out real course quiz in Slovenian, including hand-drawn circuit schematics. With state-of-the-art backends (GPT-5.2 and Gemini-3 Pro), the full pipeline achieves $\approx$8-point mean absolute difference to lecturer grades with low bias and an estimated manual-review trigger rate of $\approx$17% at $D_{\max}=40$. Ablations show that trivial prompting and removing the reference solution substantially degrade accuracy and introduce systematic over-grading, confirming that structured prompting and reference grounding are essential.

自动批改多模态大模型手写识别教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。