用视觉语言模型自动评估手写数学题,准确且能解释推理过程。
VEHME: A Vision-Language Model For Evaluating Handwritten Mathematics Expressions
- 分两阶段训练:先监督微调,再强化学习对齐评分标准。
- 在两个数据集上表现超越开源模型,接近商用系统精度。
- 引入感知表达的视觉提示模块,提升复杂布局识别能力。
自动评估手写数学解答是教育科技中的重要问题,但因格式多样、布局无序和符号复杂而极具挑战。为此,我们提出 VEHME——一个用于评估开放式手写数学表达式的视觉语言模型,具备高精度与可解释的推理轨迹。VEHME 采用两阶段训练流程:(i) 使用结构化推理数据进行监督微调;(ii) 通过强化学习对齐多维评分目标,包括正确性、推理深度和错误定位。为增强空间理解,我们提出表达感知的视觉提示模块,基于自动生成的多行数学表达式数据集训练,以在视觉异构输入中稳健引导注意力。在 AIHub 与 FERMAT 数据集上的评估显示,VEHME 在开源模型中达到顶尖性能,逼近专有系统水平,展现出作为可扩展、可访问的自动化数学评估工具的巨大潜力。训练与实验代码已公开于 GitHub。
原文摘要 · Abstract (English)
Automatically assessing handwritten mathematical solutions is an important problem in educational technology with practical applications, but it remains a significant challenge due to the diverse formats, unstructured layouts, and symbolic complexity of student work. To address this challenge, we introduce VEHME-a Vision-Language Model for Evaluating Handwritten Mathematics Expressions-designed to assess open-form handwritten math responses with high accuracy and interpretable reasoning traces. VEHME integrates a two-phase training pipeline: (i) supervised fine-tuning using structured reasoning data, and (ii) reinforcement learning that aligns model outputs with multi-dimensional grading objectives, including correctness, reasoning depth, and error localization. To enhance spatial understanding, we propose an Expression-Aware Visual Prompting Module, trained on our synthesized multi-line math expressions dataset to robustly guide attention in visually heterogeneous inputs. Evaluated on AIHub and FERMAT datasets, VEHME achieves state-of-the-art performance among open-source models and approaches the accuracy of proprietary systems, demonstrating its potential as a scalable and accessible tool for automated math assessment. Our training and experiment code is publicly available at our GitHub repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。