通过迭代聚焦提升界面元素定位精度,让AI更准点击屏幕按钮。
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
- 动态调用多个工具,逐步缩小关注区域,精准定位屏幕元素。
- 仅用1.85万样本训练,准确率达52.8%,超越多个大模型。
- 适合需要高精度界面操作的自动化测试与智能助手场景。
多模态大语言模型显著提升了图形用户界面(GUI)系统的能力,使其从受控模拟环境拓展至多样平台的真实复杂场景。然而,实际可用性仍受限于视觉定位的可靠性,即准确将文本描述映射到屏幕上的具体元素。这一限制导致系统无法精确执行点击、拖拽等指针级操作。为此,我们提出 GUI-Spotlight:一种专为图像引导推理训练的模型,能动态调用多个专用工具,迭代细化关注区域,从而显著提升视觉定位准确性。在 ScreenSpot-Pro 基准测试中,仅使用 18.5K 训练样本的 GUI-Spotlight 达到 52.8% 的准确率,优于 V2P-7B(50.6%,960万样本)和 GTA-1-7B(50.1%,156万样本)。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have markedly expanded the competence of graphical user-interface (GUI) systems, propelling them beyond controlled simulations into complex, real-world environments across diverse platforms. However, practical usefulness is still bounded by the reliability of visual grounding, i.e., mapping textual references to exact on-screen elements. This limitation prevents the system from accurately performing pointer-level actions such as clicking or dragging. To address it, we introduce GUI-Spotlight -- a model trained for image-grounded reasoning that dynamically invokes multiple specialized tools to iteratively narrow its focus to the relevant region of the screen, thereby substantially improving visual grounding accuracy. On the ScreenSpot-Pro benchmark, GUI-Spotlight trained with only 18.5K training samples achieves 52.8\% accuracy, surpassing V2P-7B (50.6\% with 9.6M training samples) and GTA-1-7B (50.1\% with 1.56M training samples).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。