arXiv:2509.17328cs.CVcs.HC2025-09ICCV被引 2

UIPro通过海量跨平台数据训练,实现通用图形界面交互能力

UIPro: Unleashing Superior Interaction Capability For GUI Agents

  • 基于2060万条界面任务数据预训练,增强对GUI的理解能力
  • 构建统一动作空间,提升跨平台任务的泛化与预测性能
  • 适合需要多平台自动化操作的研究者与开发者使用

构建能像人类一样感知和操作图形用户界面(GUI)的自主智能体,一直是人工智能领域的愿景。其核心在于GUI交互能力,涉及界面理解与规划。现有方法依赖视觉语言模型(VLMs)的多模态理解能力,但受限于场景单一、数据量不足及动作空间异构,难以实现通用化GUI代理。为此,本文提出全新通用型GUI代理UIPro,利用大规模跨平台、多任务的GUI交互数据进行训练,并引入统一动作空间。我们首先构建涵盖2060万项GUI理解任务的综合性数据集,用于预训练UIPro,使其具备强大的界面基础能力,这是下游任务的关键。随后,建立统一动作空间,整合异构任务数据集,生成合并数据集,通过持续微调提升UIPro的动作预测能力。实验结果表明,UIPro在多个平台的多种GUI任务基准测试中表现优异,验证了该方法的有效性。

原文摘要 · Abstract (English)

Building autonomous agents that perceive and operate graphical user interfaces (GUIs) like humans has long been a vision in the field of artificial intelligence. Central to these agents is the capability for GUI interaction, which involves GUI understanding and planning capabilities. Existing methods have tried developing GUI agents based on the multi-modal comprehension ability of vision-language models (VLMs). However, the limited scenario, insufficient size, and heterogeneous action spaces hinder the progress of building generalist GUI agents. To resolve these issues, this paper proposes \textbf{UIPro}, a novel generalist GUI agent trained with extensive multi-platform and multi-task GUI interaction data, coupled with a unified action space. We first curate a comprehensive dataset encompassing 20.6 million GUI understanding tasks to pre-train UIPro, granting it a strong GUI grounding capability, which is key to downstream GUI agent tasks. Subsequently, we establish a unified action space to harmonize heterogeneous GUI agent task datasets and produce a merged dataset to foster the action prediction ability of UIPro via continued fine-tuning. Experimental results demonstrate UIPro's superior performance across multiple GUI task benchmarks on various platforms, highlighting the effectiveness of our approach.

GUI代理多模态动作空间自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。