arXiv:2503.03196cs.CVcs.HC2025-03CVPR被引 21

基于视觉的GUI代理,一眼看懂界面并精准操作

SpiritSight Agent: Advanced GUI Agent with One Look

  • 构建多层级大尺度数据集,提升界面元素理解能力
  • 提出通用块解析方法,解决高分辨率输入歧义问题
  • 跨平台导航准确率领先,适合自动化交互场景

图形用户界面(GUI)代理在辅助人机交互、自动化用户操作方面展现强大能力。理想的GUI代理需具备高精度、低延迟及跨平台兼容性。近期基于视觉的方法借助先进视觉语言模型(VLMs)取得进展,虽满足兼容性和低延迟要求,但因元素定位能力有限,准确率普遍偏低。为此,我们提出视觉端到端的GUI代理SpiritSight,其在多种GUI平台上表现优异。首先,我们利用可扩展方法构建了多层级、大规模、高质量的GUI数据集GUI-Lasagne,赋予SpiritSight强大的界面理解与定位能力。其次,提出通用块解析(UBP)方法,有效解决高分辨率视觉输入中的模糊性问题,进一步增强对象定位性能。实验表明,SpiritSight在多个GUI基准测试中超越现有先进方法,展现出卓越的导航能力与跨平台兼容性。模型与数据集详见https://hzhiyuan.github.io/SpiritSight-Agent。

原文摘要 · Abstract (English)

Graphical User Interface (GUI) agents show amazing abilities in assisting human-computer interaction, automating human user's navigation on digital devices. An ideal GUI agent is expected to achieve high accuracy, low latency, and compatibility for different GUI platforms. Recent vision-based approaches have shown promise by leveraging advanced Vision Language Models (VLMs). While they generally meet the requirements of compatibility and low latency, these vision-based GUI agents tend to have low accuracy due to their limitations in element grounding. To address this issue, we propose $\textbf{SpiritSight}$, a vision-based, end-to-end GUI agent that excels in GUI navigation tasks across various GUI platforms. First, we create a multi-level, large-scale, high-quality GUI dataset called $\textbf{GUI-Lasagne}$ using scalable methods, empowering SpiritSight with robust GUI understanding and grounding capabilities. Second, we introduce the $\textbf{Universal Block Parsing (UBP)}$ method to resolve the ambiguity problem in dynamic high-resolution of visual inputs, further enhancing SpiritSight's ability to ground GUI objects. Through these efforts, SpiritSight agent outperforms other advanced methods on diverse GUI benchmarks, demonstrating its superior capability and compatibility in GUI navigation tasks. Models and datasets are available at https://hzhiyuan.github.io/SpiritSight-Agent.

GUI代理视觉语言模型界面理解自动化交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。