评测视觉语言模型在学生手写数学题上的理解能力
DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images
- 构建包含2030张学生手写数学题的标注数据集
- 顶尖模型在真实教育场景下仍表现不足
- 可用自动生成问题评估模型,节省人工成本
在真实教育场景中,视觉语言模型需处理自然、杂乱的图像以及特定领域的语言与概念。为评估模型在类似情境下的潜力,我们提出DrawEduMath,一个包含2,030张学生手写数学题的英文数据集。教师提供了详尽标注,包括每张图的自由描述及11,661个问答对,涵盖解题策略、作图结构与书写特征等教学洞察。我们在教师撰写的问答对上评估模型,同时使用语言模型基于教师描述生成44,362条合成问答对进行对比。结果表明,即使最先进的视觉语言模型在该数据集上仍有较大提升空间;而合成问答对虽不完美,但能给出与真人标注相似的模型排名。我们公开DrawEduMath,以支持面向教育场景的视觉语言模型数学推理能力评估。
原文摘要 · Abstract (English)
In real-world settings, vision language models (VLMs) should robustly handle naturalistic, noisy visual content as well as domain-specific language and concepts. For example, K-12 educators using digital learning platforms may need to examine and provide feedback across many images of students' math work. To assess the potential of VLMs to support educators in settings like this one, we introduce DrawEduMath, an English-language dataset of 2,030 images of students' handwritten responses to K-12 math problems. Teachers provided detailed annotations, including free-form descriptions of each image and 11,661 question-answer (QA) pairs. These annotations capture a wealth of pedagogical insights, ranging from students' problem-solving strategies to the composition of their drawings, diagrams, and writing. We evaluate VLMs on teachers' QA pairs, as well as 44,362 synthetic QA pairs derived from teachers' descriptions using language models (LMs). We show that even state-of-the-art VLMs leave much room for improvement on DrawEduMath questions. We also find that synthetic QAs, though imperfect, can yield similar model rankings as teacher-written QAs. We release DrawEduMath to support the evaluation of VLMs' abilities to reason mathematically over images gathered with educational contexts in mind.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。