测试视觉语言模型对手写数学题的判错与纠错能力
Can Vision-Language Models Evaluate Handwritten Math?

- 构建了涵盖4类错误的2200+手写数学题基准集
- 现有模型纠错率最高仅77%,手写识别能力普遍较弱
- 适合研究AI自动批改、教育科技与多模态模型改进者
视觉语言模型(VLMs)在自动批改手写学生作答方面展现出新潜力,尤其在数学领域。然而,尚无系统性研究评估VLMs对手写内容的检测、定位和纠错能力。为此,我们提出FERMAT基准,用于评估VLMs在手写数学内容中的错误识别与修正能力。FERMAT覆盖计算、概念、符号和呈现四类错误,包含609道人工精心设计的7-12年级题目,共生成超过2200份带有故意扰动的手写解答。我们使用FERMAT对九种VLMs在三个任务(错误检测、定位、纠正)上进行评测。结果显示,当前VLMs在手写文本推理方面存在显著不足,其中Gemini-1.5-Pro纠错率最高达77%。部分模型在输入为印刷体或图像时表现明显提升,表明其对手写内容处理能力有限。这些发现揭示了现有模型的局限性,并指明改进方向。我们已开源FERMAT及全部资源,以推动后续研究。
原文摘要 · Abstract (English)
Recent advancements in Vision-Language Models (VLMs) have opened new possibilities in automatic grading of handwritten student responses, particularly in mathematics. However, a comprehensive study to test the ability of VLMs to evaluate and reason over handwritten content remains absent. To address this gap, we introduce FERMAT, a benchmark designed to assess the ability of VLMs to detect, localize and correct errors in handwritten mathematical content. FERMAT spans four key error dimensions - computational, conceptual, notational, and presentation - and comprises over 2,200 handwritten math solutions derived from 609 manually curated problems from grades 7-12 with intentionally introduced perturbations. Using FERMAT we benchmark nine VLMs across three tasks: error detection, localization, and correction. Our results reveal significant shortcomings in current VLMs in reasoning over handwritten text, with Gemini-1.5-Pro achieving the highest error correction rate (77%). We also observed that some models struggle with processing handwritten content, as their accuracy improves when handwritten inputs are replaced with printed text or images. These findings highlight the limitations of current VLMs and reveal new avenues for improvement. We release FERMAT and all the associated resources in the open-source to drive further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。