arXiv:2502.01977cs.CV2025-02ACL被引 24

用大模型自动标注界面功能,规模达70万条,提升软件自动化能力

AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs

  • 通过对比模拟操作前后的界面变化,用大模型推断元素功能
  • 构建了70.4万条高质量功能标注数据集,准确率接近人工水平
  • 适用于界面理解、智能助手等任务,适合做UI自动化研究者

视觉语言模型在用户界面理解方面备受关注,因其在提升软件自动化方面的潜力。然而,现有数据集或仅提供大规模无上下文的元素标注,或仅在小规模下提供带上下文的功能描述。本文提出 extbf{AutoGUI} 流水线,可规模化自动为界面元素添加详细功能描述。具体而言,利用大语言模型(LLMs)通过比较模拟交互前后界面状态的变化来推断元素功能。为提升标注质量,提出基于LLM的拒绝与验证机制,无需人工干预即可剔除无效标注。我们使用该流程构建了高质量的 AutoGUI-704k 数据集,包含多样且详尽的功能标注,远超以往数据集。人类评估显示,其标注正确率可媲美训练过的专业标注员。大量实验表明,该数据集显著增强视觉语言模型的界面定位能力,并展现出显著的扩展效应。此外,我们还展示了该数据集在界面智能体任务中的潜在应用价值。

原文摘要 · Abstract (English)

User interface understanding with vision-language models (VLMs) has received much attention due to its potential for enhancing software automation. However, existing datasets used to build UI-VLMs either only contain large-scale context-free element annotations or contextualized functional descriptions for elements at a small scale. In this work, we propose the \textbf{AutoGUI} pipeline for automatically annotating UI elements with detailed functionality descriptions at scale. Specifically, we leverage large language models (LLMs) to infer element functionality by comparing UI state changes before and after simulated interactions. To improve annotation quality, we propose LLM-aided rejection and verification, eliminating invalid annotations without human labor. We construct a high-quality AutoGUI-704k dataset using the proposed pipeline, featuring diverse and detailed functionality annotations that are hardly provided by previous datasets. Human evaluation shows that we achieve annotation correctness comparable to a trained human annotator. Extensive experiments show that our dataset remarkably enhances VLM's UI grounding capabilities and exhibits significant scaling effects. We also show the interesting potential use of our dataset in UI agent tasks. Please view our project at https://autogui-project.github.io/.

界面理解大模型自动化数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。