arXiv:2601.06944cs.CVcs.AI2026-01

用手绘图诊断题评测大模型,发现当前模型在符号化和噪声场景下表现远不及人类。

SketchJudge: A Diagnostic Benchmark for Grading Hand-drawn Diagrams with Multimodal Large Language Models

  • 构建手绘STEM图评分基准SketchJudge,覆盖四大学科领域
  • 1015份学生手绘作答中含多样风格与错误类型,揭示模型缺陷
  • 实测表明先进大模型评分能力显著落后人类,凸显视觉语言对齐不足

尽管多模态大语言模型在视觉理解方面取得显著进展,但在面对人类绘制的非结构化、模糊草图时仍表现不佳。这一局限在视觉评分任务中尤为突出,要求模型不仅解决问题,还需诊断手绘图中的错误。此类诊断依赖于复杂的结构、语义与元认知推理。为此,我们提出SketchJudge,一个专为评估多模态大模型作为手绘图评分器而设计的新基准。该基准包含1,015份跨几何、物理、图表和流程图四个领域的手绘学生作答,涵盖多种风格变异和不同类型的错误。在SketchJudge上的评估表明,即使最先进的多模态大模型也显著落后于人类,验证了该基准在揭示当前视觉-语言对齐在符号化和噪声环境下脆弱性方面的有效性。所有数据、代码和评估脚本已公开于https://github.com/yuhangsu82/SketchJudge。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual understanding, they often struggle when faced with the unstructured and ambiguous nature of human-generated sketches. This limitation is particularly pronounced in the underexplored task of visual grading, where models should not only solve a problem but also diagnose errors in hand-drawn diagrams. Such diagnostic capabilities depend on complex structural, semantic, and metacognitive reasoning. To bridge this gap, we introduce SketchJudge, a novel benchmark tailored for evaluating MLLMs as graders of hand-drawn STEM diagrams. SketchJudge encompasses 1,015 hand-drawn student responses across four domains: geometry, physics, charts, and flowcharts, featuring diverse stylistic variations and distinct error types. Evaluations on SketchJudge demonstrate that even advanced MLLMs lag significantly behind humans, validating the benchmark's effectiveness in exposing the fragility of current vision-language alignment in symbolic and noisy contexts. All data, code, and evaluation scripts are publicly available at https://github.com/yuhangsu82/SketchJudge.

多模态手绘识别评分基准教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。