评测大模型读不懂学生手写数学错题,提出新基准提升教育AI诊断能力
Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math
- 构建手写数学错题分析基准ScratchMath,含1720个真实样本
- 模型在视觉识别与逻辑推理上表现远低于人类专家,尤其开源模型差距大
- 适合教育AI、多模态模型研究者,推动个性化学习反馈发展
评估学生手写演算过程对个性化教育反馈至关重要,但受笔迹多样、布局复杂及解题方式差异影响,挑战重重。现有教育NLP多关注文本回答,忽视手写演算的多模态特性。当前多模态大模型(MLLMs)虽擅长视觉推理,但常以「应试者视角」为主,侧重输出正确答案而非诊断错误。为此,我们提出ScratchMath,一个专为解释和分类真实手写数学错误设计的新基准。数据集包含1720份中国中小学学生手写样本,支持两大任务:错误原因解释(ECE)与错误类型分类(ECC),定义七类错误。通过多轮专家标注、审核与验证,实现严谨标注。我们系统评估了16个主流MLLMs,发现其性能显著低于人类专家,尤其在视觉识别与逻辑推理方面。专有模型明显优于开源模型,大型推理模型在错误解释中展现潜力。所有评估数据与框架已公开,供后续研究使用。
原文摘要 · Abstract (English)
Assessing student handwritten scratchwork is crucial for personalized educational feedback but presents unique challenges due to diverse handwriting, complex layouts, and varied problem-solving approaches. Existing educational NLP primarily focuses on textual responses and neglects the complexity and multimodality inherent in authentic handwritten scratchwork. Current multimodal large language models (MLLMs) excel at visual reasoning but typically adopt an "examinee perspective", prioritizing generating correct answers rather than diagnosing student errors. To bridge these gaps, we introduce ScratchMath, a novel benchmark specifically designed for explaining and classifying errors in authentic handwritten mathematics scratchwork. Our dataset comprises 1,720 mathematics samples from Chinese primary and middle school students, supporting two key tasks: Error Cause Explanation (ECE) and Error Cause Classification (ECC), with seven defined error types. The dataset is meticulously annotated through rigorous human-machine collaborative approaches involving multiple stages of expert labeling, review, and verification. We systematically evaluate 16 leading MLLMs on ScratchMath, revealing significant performance gaps relative to human experts, especially in visual recognition and logical reasoning. Proprietary models notably outperform open-source models, with large reasoning models showing strong potential for error explanation. All evaluation data and frameworks are publicly available to facilitate further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。