arXiv:2603.00925cs.CLcs.CV2026-03

11个视觉语言模型在学生错题上表现差,误判率高。

The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors

  • 测试11个视觉语言模型对真实学生手写数学作答的分析能力。
  • 模型在低水平学生作答上准确率下降,尤其错题评估错误率超60%。
  • 适合教育领域研究者关注模型在教学诊断中的可靠性问题。

有效的数学教育需要识别并回应学生的错误。为使AI支持教学应用,模型必须在不同学生能力水平下表现良好。我们的工作提供了对11个视觉语言模型(VLMs)在DrawEduMath基准上的全年性能快照,该基准包含真实学生的手写、手绘数学答题。结果发现,所有评估的VLMs在描述需要更多教学帮助的学生作答时均表现不佳,且在所有问答任务中,对评估学生错误的问题表现最差。尽管这些模型可能被优化为解题专家,但我们的结果表明,它们需新的开发激励机制才能充分支持教育应用场景。

原文摘要 · Abstract (English)

Effective mathematics education requires identifying and responding to students' mistakes. For AI to support pedagogical applications, models must perform well across different levels of student proficiency. Our work provides an extensive, year-long snapshot of how 11 vision-language models (VLMs) perform on DrawEduMath, a QA benchmark involving real students' handwritten, hand-drawn responses to math problems. We find that models' weaknesses concentrate on a core component of math education: student error. All evaluated VLMs underperform when describing work from students who require more pedagogical help, and across all QA, they struggle the most on questions related to assessing student error. Thus, while VLMs may be optimized to be math problem solving experts, our results suggest that they require alternative development incentives to adequately support educational use cases.

教育AI视觉语言模型错题诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。