用缩放操作提升界面元素定位精度,无需训练即可生效。
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
- 利用缩放的四个特性实现动态聚焦与上下文自适应
- 在ScreenSpot-Pro上使UI-Venus-72B成功率达73.1%,领先现有模型
- 适合做界面智能代理的开发者或研究者参考
界面元素定位是构建图形用户界面(GUI)智能体的基础能力。尽管现有方法依赖大规模边界框标注,仍面临跨平台泛化、复杂布局分析和细粒度定位等挑战。本文探索缩放作为强但未被充分挖掘的先验信息,提出无需训练的ZoomClick方法。通过刻画缩放的四个关键属性(预缩放状态、缩放深度、缩小尺寸、最小裁剪尺寸),充分释放其在动态空间聚焦与自适应上下文切换中的潜力。实验表明,该方法显著提升通用视觉语言模型及专用GUI定位模型性能,在多个主流基准上达到最先进水平;例如,UI-Venus-72B在ScreenSpot-Pro上取得73.1%的成功率。此外,我们构建GUIZoom-Bench基准,用于评估模型对缩放变化的适应能力,旨在推动未来在训练与测试阶段利用缩放进行扩展的研究。
原文摘要 · Abstract (English)
Grounding is a fundamental capability for building graphical user interface (GUI) agents. Although existing approaches rely on large-scale bounding box supervision, they still face various challenges, such as cross-platform generalization, complex layout analysis, and fine-grained element localization. In this paper, we investigate zoom as a strong yet underexplored prior for GUI grounding, and propose a training-free method, ZoomClick. By characterizing four key properties of zoom (i.e., pre-zoom, depth, shrink size, minimal crop size), we unlock its full capabilities for dynamic spatial focusing and adaptive context switching. Experiments demonstrate that our method significantly boosts the performance of both general vision-language and specialized GUI grounding models, achieving state-of-the-art results on several mainstream benchmarks; for example, UI-Venus-72B attains a 73.1% success rate on ScreenSpot-Pro. Furthermore, we present GUIZoom-Bench, a benchmark for evaluating model adaptability to zoom, aiming to inspire future research on improving zoom for further training and test-time scaling in GUI grounding tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。