提出GUI知识基准,揭示视觉语言模型在界面任务中的认知短板。
GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
- 从界面、交互、流程三方面提炼GUI核心知识维度
- 在6大平台292个应用上构建多选与判断题基准测试
- 发现模型懂组件功能但不懂状态跟踪与任务进度评估
视觉语言模型(VLMs)在图形用户界面(GUI)自动化任务中取得进展,但仍远落后于人类。我们推测这一差距源于缺乏核心的GUI知识,而现有训练方法(如监督微调和强化学习)难以完全弥补。通过分析常见任务失败模式,我们将GUI知识提炼为三个维度:(1) 界面知识——组件功能、布局语义与系统状态;(2) 交互知识——交互类型及其影响;(3) 流程知识——任务目标与操作序列。为此,我们构建了GUI Knowledge Bench,涵盖六个平台(Web、Android、MacOS、Windows、Linux、iOS)和292个应用,包含多项选择与是非题。评估显示,当前VLMs普遍了解单个组件的功能,但缺乏追踪系统状态、遵循交互规范及判断任务完成进度的特定知识。真实场景任务实验进一步验证了GUI知识与任务成功率之间的紧密关联。本工作为评估VLM的GUI能力提供结构化框架,支持下游训练前筛选更优模型,并为构建更强的GUI智能体提供方向。
原文摘要 · Abstract (English)
Vision language models (VLMs) have advanced graphical user interface (GUI) task automation but still lag behind humans. We hypothesize this gap stems from missing core GUI knowledge, which existing training schemes (such as supervised fine tuning and reinforcement learning) alone cannot fully address. By analyzing common failure patterns in GUI task execution, we distill GUI knowledge into three dimensions: (1) interface knowledge about widget functions, layout semantics, and system states; (2) interaction knowledge about GUI interaction types and effects; and (3) procedure knowledge of task objectives and workflow sequences. We further introduce GUI Knowledge Bench, a benchmark with multiple-choice and yes/no questions across six platforms (Web, Android, MacOS, Windows, Linux, IOS) and 292 applications. Our evaluation indicates that current VLMs are generally aware of the functions of individual widgets, but lack the GUI-specific knowledge required to track system states, adhere to GUI interaction conventions, and assess task completion progress. Experiments on real-world GUI tasks further validate the close link between GUI knowledge and task success. By providing a structured framework for assessing GUI knowledge, our work supports the selection of VLMs with greater potential prior to downstream training and provides insights for building more capable GUI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。