评测大模型在小学真实视觉题上的推理能力,发现其空间推理有明显短板。
Visual Reasoning Benchmark: Evaluating Multimodal LLMs on Classroom-Authentic Visual Problems from Primary Education
- 构建来自赞比亚和印度小学的701道原生视觉题数据集
- 模型在计数等静态任务表现好,但折叠、旋转等动态操作能力差
- 适合教育AI评估,帮助识别模型在课堂应用中的潜在风险
尽管人工智能模型在文本推理上已达到顶尖水平,但在空间与关系结构推理方面仍存在关键瓶颈,尤其是在依赖视觉的小学数学中。本文提出视觉推理基准(VRB),一个针对多模态大语言模型(MLLM)的真实课堂视觉问题评估数据集。该数据集包含701道来自赞比亚和印度小学考试的题目,涵盖类比推理、模式补全、空间匹配等任务。基准设计强调使用未经编辑、文字极少的图像,以测试模型是否能满足小学教育的实际需求。结果显示,模型在计数、缩放等静态技能上表现良好,但在折叠、反射、旋转等动态操作上存在明显“空间天花板”。这些弱点可能导致课堂使用中误判、错误引导甚至强化学生误解。因此,像VRB这样的教育导向基准对界定多模态工具在教学中的实际边界至关重要。
原文摘要 · Abstract (English)
AI models have achieved state-of-the-art results in textual reasoning; however, their ability to reason over spatial and relational structures remains a critical bottleneck -- particularly in early-grade maths, which relies heavily on visuals. This paper introduces the visual reasoning benchmark (VRB), a novel dataset designed to evaluate Multimodal Large Language Models (MLLMs) on their ability to solve authentic visual problems from classrooms. This benchmark is built on a set of 701 questions sourced from primary school examinations in Zambia and India, which cover a range of tasks such as reasoning by analogy, pattern completion, and spatial matching. We outline the methodology and development of the benchmark which intentionally uses unedited, minimal-text images to test if models can meet realistic needs of primary education. Our findings reveal a ``jagged frontier'' of capability where models demonstrate better proficiency in static skills such as counting and scaling, but reach a distinct ``spatial ceiling'' when faced with dynamic operations like folding, reflection, and rotation. These weaknesses pose a risk for classroom use on visual reasoning problems, with the potential for incorrect marking, false scaffolding, and reinforcing student misconceptions. Consequently, education-focused benchmarks like the VRB are essential for determining the functional boundaries of multimodal tools used in classrooms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。