现有AI模型常误改学生手写数学题,新方法能识别并惩罚这种过度修正。
When VLMs 'Fix' Students: Identifying and Penalizing Over-Correction in the Evaluation of Multi-line Handwritten Math OCR

- 用大语言模型结合评分标准,设计新评估指标PINK
- 在FERMAT数据集上15个模型排名反转,GPT-4o因纠错过度被重罚
- 人类专家更认可新方法,正确率比传统方法高15.5个百分点
手写数学准确转录对教育AI至关重要,但现有基准无法有效评估多行解题过程。多数研究聚焦单行表达式,依赖BLEU等词法指标,难以衡量跨行语义推理。本文首次系统研究多行手写数学OCR,揭示视觉语言模型(VLMs)的关键缺陷:过度纠正。模型常主动修正学生错误,掩盖真实问题。为此提出PINK(基于墨水的惩罚评分),利用大语言模型进行基于评分标准的评分,并明确惩罚过度纠正行为。在FERMAT数据集上对15个前沿VLM的全面评估显示,与BLEU相比排名显著反转:GPT-4o因激进纠错被重罚,Gemini 2.5 Flash成为最忠实转录者。人工专家实验表明,PINK与人类判断一致性更高(55.0%偏好,优于BLEU的39.5%),为教育场景下的手写数学OCR提供更可靠评估框架。
原文摘要 · Abstract (English)
Accurate transcription of handwritten mathematics is crucial for educational AI systems, yet current benchmarks fail to evaluate this capability properly. Most prior studies focus on single-line expressions and rely on lexical metrics such as BLEU, which fail to assess the semantic reasoning across multi-line student solutions. In this paper, we present the first systematic study of multi-line handwritten math Optical Character Recognition (OCR), revealing a critical failure mode of Vision-Language Models (VLMs): over-correction. Instead of faithfully transcribing a student's work, these models often "fix" errors, thereby hiding the very mistakes an educational assessment aims to detect. To address this, we propose PINK (Penalized INK-based score), a semantic evaluation metric that leverages a Large Language Model (LLM) for rubric-based grading and explicitly penalizes over-correction. Our comprehensive evaluation of 15 state-of-the-art VLMs on the FERMAT dataset reveals substantial ranking reversals compared to BLEU: models like GPT-4o are heavily penalized for aggressive over-correction, whereas Gemini 2.5 Flash emerges as the most faithful transcriber. Furthermore, human expert studies show that PINK aligns significantly better with human judgment (55.0% preference over BLEU's 39.5%), providing a more reliable evaluation framework for handwritten math OCR in educational settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。