arXiv:2411.18932cs.CLcs.AI2024-11NAACL被引 6

用儿童编程语言测试大模型的综合逻辑能力

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges

  • 基于Scratch构建多模态编程评测基准,融合视觉与代码逻辑
  • 模型需理解图形与程序结构,实现统一逻辑推理
  • 适合关注多模态推理与教育类应用的研究者

近年来,大型多模态模型(LMMs)在代码生成方面表现出色,主要通过图像到代码的基准进行评估。然而,这些基准仅限于特定的可视化编程场景,将逻辑推理与多模态理解能力割裂。为此,我们提出ScratchEval,一个新型基准,用于评估LMMs在可视化编程中的推理能力。ScratchEval基于广泛应用于儿童编程教育的积木式编程语言Scratch,通过整合视觉元素与嵌入式编程逻辑,要求模型同时处理视觉信息与代码结构,从而全面评估其对编程意图的理解能力。我们的评估方法超越传统的图像到代码映射,聚焦于统一的逻辑思维与问题解决能力,为评估LMMs在可视化编程中的表现提供更全面、更具挑战性的框架。ScratchEval不仅弥补了现有评估方法的不足,也为LMMs在该领域的未来发展提供了新视角。基准可访问 https://github.com/HKBUNLP/ScratchEval。

原文摘要 · Abstract (English)

Recent advancements in large multimodal models (LMMs) have showcased impressive code generation capabilities, primarily evaluated through image-to-code benchmarks. However, these benchmarks are limited to specific visual programming scenarios where the logic reasoning and the multimodal understanding capacities are split apart. To fill this gap, we propose ScratchEval, a novel benchmark designed to evaluate the visual programming reasoning ability of LMMs. ScratchEval is based on Scratch, a block-based visual programming language widely used in children's programming education. By integrating visual elements and embedded programming logic, ScratchEval requires the model to process both visual information and code structure, thereby comprehensively evaluating its programming intent understanding ability. Our evaluation approach goes beyond the traditional image-to-code mapping and focuses on unified logical thinking and problem-solving abilities, providing a more comprehensive and challenging framework for evaluating the visual programming ability of LMMs. ScratchEval not only fills the gap in existing evaluation methods, but also provides new insights for the future development of LMMs in the field of visual programming. Our benchmark can be accessed at https://github.com/HKBUNLP/ScratchEval .

多模态模型编程评估教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。