arXiv:2507.23779cs.CVcs.AI2025-07被引 21

提升图形界面理解能力,让智能助手更准地点击和输入。

Phi-Ground Tech Report: Advancing Perception in GUI Grounding

  • 通过系统性实验优化数据与训练细节,构建高效视觉-交互模型。
  • 在ScreenSpot-pro和UI-Vision上分别达到43.2和27.2的领先准确率。
  • 适合研究智能代理、人机交互及多模态感知的开发者参考。

随着多模态推理模型的发展,类贾维斯的计算机使用代理(CUAs)正成为现实。图形界面(GUI) grounding 是其实现真实操作的核心,类似于机器人中的机械控制,直接决定系统成败。它决定了点击、输入等动作及其坐标参数。当前端到端接地模型在ScreenSpot-pro和UI-Vision等挑战性基准上的准确率仍低于65%,远未达到部署要求。本文开展了一项关于接地模型训练的实证研究,涵盖从数据收集到模型训练的各个环节。最终我们提出了Phi-Ground模型系列,在所有五个接地基准上均取得参数量小于100亿的模型中最佳表现。在端到端设置下,模型在ScreenSpot-pro上达到43.2分,在UI-Vision上达到27.2分,显著优于现有方法。我们认为本文探讨的各项细节,以及成功与失败经验,不仅有助于构建更可靠的接地模型,也对其他感知任务具有参考价值。

原文摘要 · Abstract (English)

With the development of multimodal reasoning models, Computer Use Agents (CUAs), akin to Jarvis from \textit{"Iron Man"}, are becoming a reality. GUI grounding is a core component for CUAs to execute actual actions, similar to mechanical control in robotics, and it directly leads to the success or failure of the system. It determines actions such as clicking and typing, as well as related parameters like the coordinates for clicks. Current end-to-end grounding models still achieve less than 65\% accuracy on challenging benchmarks like ScreenSpot-pro and UI-Vision, indicating they are far from being ready for deployment. % , as a single misclick can result in unacceptable consequences. In this work, we conduct an empirical study on the training of grounding models, examining details from data collection to model training. Ultimately, we developed the \textbf{Phi-Ground} model family, which achieves state-of-the-art performance across all five grounding benchmarks for models under $10B$ parameters in agent settings. In the end-to-end model setting, our model still achieves SOTA results with scores of \textit{\textbf{43.2}} on ScreenSpot-pro and \textit{\textbf{27.2}} on UI-Vision. We believe that the various details discussed in this paper, along with our successes and failures, not only clarify the construction of grounding models but also benefit other perception tasks. Project homepage: \href{https://zhangmiaosen2000.github.io/Phi-Ground/}{https://zhangmiaosen2000.github.io/Phi-Ground/}

GUI理解智能代理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。