arXiv:2602.10814cs.AI2026-02被引 1

评测AI在图形化编程中看、想、操作的能力,发现其规划强但动手弱。

See, Plan, Snap: Evaluating Multimodal GUI Agents in Scratch

  • 设计新基准ScratchWorld,分四类任务评估AI构建程序能力
  • 两种交互模式分离判断:精细拖拽与高级语义指令,定位失败原因
  • 实测发现顶尖模型有强计划力但细粒度操作差,适合教育类AI研究者

基于积木式编程环境的低代码教育中,评估AI代理通过图形用户界面(GUI)构建程序的能力仍不充分。本文提出ScratchWorld,一个针对多模态GUI代理在Scratch中进行程序构建任务的评测基准。该基准以‘使用-修改-创造’教学框架为基础,包含83个精心设计的任务,涵盖创建、调试、扩展和计算四类问题。为精准诊断代理失败原因,基准采用两种互补交互模式:原始模式要求精细拖拽操作以直接评估视觉运动控制能力;复合模式使用高层语义API,将程序推理与GUI执行解耦。为确保评估可靠性,提出基于执行的评测协议,通过浏览器环境中的运行测试验证生成程序的功能正确性。对前沿多模态语言模型与GUI代理的广泛实验揭示显著的‘推理-执行’差距,表明尽管规划能力较强,但在细粒度GUI操作方面仍存在持续挑战。

原文摘要 · Abstract (English)

Block-based programming environments such as Scratch play a central role in low-code education, yet evaluating the capabilities of AI agents to construct programs through Graphical User Interfaces (GUIs) remains underexplored. We introduce ScratchWorld, a benchmark for evaluating multimodal GUI agents on program-by-construction tasks in Scratch. Grounded in the Use-Modify-Create pedagogical framework, ScratchWorld comprises 83 curated tasks spanning four distinct problem categories: Create, Debug, Extend, and Compute. To rigorously diagnose the source of agent failures, the benchmark employs two complementary interaction modes: primitive mode requires fine-grained drag-and-drop manipulation to directly assess visuomotor control, while composite mode uses high-level semantic APIs to disentangle program reasoning from GUI execution. To ensure reliable assessment, we propose an execution-based evaluation protocol that validates the functional correctness of the constructed Scratch programs through runtime tests within the browser environment. Extensive experiments across state-of-the-art multimodal language models and GUI agents reveal a substantial reasoning--acting gap, highlighting persistent challenges in fine-grained GUI manipulation despite strong planning capabilities.

GUI代理编程教育多模态评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。