构建首个评估视觉语言模型操作设计软件能力的基准
CANVAS: A Benchmark for Vision-Language Models on Tool-Based User Interface Design
- 设计任务分复制与修改两类,通过工具调用逐步更新界面
- 涵盖598个任务,基于30类功能的3.3K个真实移动端界面
- 揭示主流模型策略性工具调用能力,指导未来优化方向
用户界面(UI)设计是设计师使用Figma或Sketch等工具反复迭代的过程。近年来,具备工具调用能力的视觉语言模型(VLMs)展现出通过迭代操作设计软件编辑界面的潜力。然而,由于缺乏针对工具化设计能力的评测基准,该能力水平尚不明确。为此,我们提出CANVAS——一个面向工具化用户界面设计的VLM基准。该基准包含598个工具调用任务,配对来自30种功能类别(如引导页、消息)的3.3K个真实移动端界面。每项任务中,VLM通过上下文相关的工具调用(如创建矩形作为按钮背景)逐步更新设计,模拟真实设计软件操作。CANVAS包含两种任务类型:(i) 设计复现,评估完整界面还原能力;(ii) 设计修改,评估对现有界面特定部分的调整能力。结果表明,领先模型表现出更策略性的工具调用,提升了设计质量。同时,我们识别出模型常见错误模式,为未来提升工具化设计能力提供方向。
原文摘要 · Abstract (English)
User interface (UI) design is an iterative process in which designers progressively refine their work with design software such as Figma or Sketch. Recent advances in vision language models (VLMs) with tool invocation suggest these models can operate design software to edit a UI design through iteration. Understanding and enhancing this capacity is important, as it highlights VLMs' potential to collaborate with designers within conventional software. However, as no existing benchmark evaluates tool-based design performance, the capacity remains unknown. To address this, we introduce CANVAS, a benchmark for VLMs on tool-based user interface design. Our benchmark contains 598 tool-based design tasks paired with ground-truth references sampled from 3.3K mobile UI designs across 30 function-based categories (e.g., onboarding, messaging). In each task, a VLM updates the design step-by-step through context-based tool invocations (e.g., create a rectangle as a button background), linked to design software. Specifically, CANVAS incorporates two task types: (i) design replication evaluates the ability to reproduce a whole UI screen; (ii) design modification evaluates the ability to modify a specific part of an existing screen. Results suggest that leading models exhibit more strategic tool invocations, improving design quality. Furthermore, we identify common error patterns models exhibit, guiding future work in enhancing tool-based design capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。