构建可控制的GUI生成评估基准,解决标注混乱与场景覆盖不足问题
SchemaGUI: A Schema-Driven Benchmark for Controllable GUI Generation Evaluation

- 基于参数化界面模板生成带确定性标注的任务
- 27B模型几何控制准确率仅75.3%,布局复杂时坐标漂移严重
- 适合研究可控生成、评估基准设计的开发者和研究人员
大型语言模型在图形用户界面生成方面展现出巨大潜力,但可靠评估仍面临数据分布不可控、标注噪声大及布局场景覆盖有限等挑战。为此,我们提出SchemaGUI,一个基于模板的可控GUI生成评估基准。通过从参数化界面模板中合成成对的自然语言指令与确定性函数调用引用,SchemaGUI可在秒级生成数千个具有确定性标注的任务,无需人工标注。基于每种场景与语言下1,000个评估实例,我们在六种代表性双语场景中对五款主流模型(包括Qwen3.5系列、Qwen3-Coder-30B和DeepSeek-R1)进行了评测。深入分析揭示三大关键发现:首先,精确的几何空间控制仍是主要瓶颈;尽管将Qwen3.5从4B扩展到27B使界面可行性从91.56%提升至99.63%,但几何得分仅从67.05%提高到75.30%。其次,生成难度高度依赖布局复杂度,当前模型在简单顺序排列上表现良好,但在密集网格和多区域组合中出现严重坐标漂移。第三,思维模式虽增加令牌消耗,但通常降低界面得分,尤其对小型模型影响更显著。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated strong potential in graphical user interface (GUI) generation, but reliable evaluation remains challenging due to uncontrolled data distributions, noisy annotations, and limited layout scenario coverage. To address this, we propose SchemaGUI, a template-based benchmark for controllable GUI generation evaluation. By synthesizing paired natural language instructions and deterministic function-call references from parameterized interface schemas, SchemaGUI can generate thousands of deterministically annotated tasks in seconds without human labeling. Based on 1,000 evaluated instances per scenario and language across six representative bilingual scenarios, we benchmark five mainstream models, including the Qwen3.5 family, Qwen3-Coder-30B, and DeepSeek-R1. Our extensive analysis reveals three key insights. First, precise geometric spatial control remains an important bottleneck; while scaling Qwen3.5 from 4B to 27B improves Schema Feasibility from 91.56% to 99.63%, the Geometry score improves more modestly (from 67.05% to 75.30%). Second, generation difficulty is highly sensitive to layout complexity, with current LLMs excelling at simple sequential arrangements but suffering severe coordinate drift in dense grids and multi-region compositions. Third, thinking mode increases token consumption while generally reducing GUI Score, particularly for smaller models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。