arXiv:2504.11257cs.HCcs.CL2025-04ACL被引 20

用AI自动生成海量GUI指令数据,提升智能助手理解界面的能力

UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis

  • 用GPT-4o自动合成复杂GUI指令数据,替代人工标注
  • 在新基准上实现领先性能,尤其在小元素和隐式指令上表现优异
  • 适合研究GUI智能体、人机交互与自动化工具的开发者

大型视觉语言模型正推动基于视觉的图形用户界面(GUI)智能体发展,这类方法相比依赖平台元数据的方案更具通用性。然而,将用户指令映射到屏幕元素位置的GUI指令定位仍是关键挑战,主要受限于公开训练数据稀缺和人工标注成本高。本文针对元素占比差异、类别分布不均及隐式指令等问题,提出大规模数据合成管道UI-E2I-Synth,利用GPT-4o生成多样化复杂指令数据。同时构建新基准UI-I2E-Bench,涵盖多维度标注,弥补现有评估局限。基于合成数据训练的模型在指令定位任务中表现优异,验证了数据合成的有效性。该基准附带详尽分析,为未来研究提供实用参考。相关资源将在https://microsoft.github.io/FIVE-UI-Evol/ 公开。

原文摘要 · Abstract (English)

Recent advancements in Large Vision-Language Models are accelerating the development of Graphical User Interface (GUI) agents that utilize human-like vision perception capabilities to enhance productivity on digital devices. Compared to approaches predicated on GUI metadata, which are platform-dependent and vulnerable to implementation variations, vision-based approaches offer broader applicability. In this vision-based paradigm, the GUI instruction grounding, which maps user instruction to the location of corresponding element on the given screenshot, remains a critical challenge, particularly due to limited public training dataset and resource-intensive manual instruction data annotation. In this paper, we delve into unexplored challenges in this task including element-to-screen ratio, unbalanced element type, and implicit instruction. To address these challenges, we introduce a large-scale data synthesis pipeline UI-E2I-Synth for generating varying complex instruction datasets using GPT-4o instead of human annotators. Furthermore, we propose a new GUI instruction grounding benchmark UI-I2E-Bench, which is designed to address the limitations of existing benchmarks by incorporating diverse annotation aspects. Our model, trained on the synthesized data, achieves superior performance in GUI instruction grounding, demonstrating the advancements of proposed data synthesis pipeline. The proposed benchmark, accompanied by extensive analyses, provides practical insights for future research in GUI grounding. We will release corresponding artifacts at https://microsoft.github.io/FIVE-UI-Evol/ .

GUI智能体指令生成大模型应用自动化测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。