用快慢双系统提升复杂界面操作的准确率
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
- 模仿人类思维模式,分快速判断与深度分析两阶段处理界面指令
- 在ScreenSpot上达77.4%准确率,在ScreenSpot-Pro上提升13.3%
- 适合需要理解嵌套层级关系的复杂界面交互任务
人类能根据任务复杂度灵活切换思考模式:从快速直觉判断到深入分析。然而,现有图形用户界面(GUI)定位系统仅依赖即时预测,缺乏推理能力,难以理解具有嵌套结构和层级关系的复杂界面布局,限制了其在复杂场景下的表现。受人类双系统认知启发,我们提出Focus框架,通过动态适应任务复杂度,在快速与深度处理间切换,兼顾效率与精度。该框架将界面定位分解为三阶段:界面摘要、视觉聚焦分析与精确坐标预测,实现对界面布局与视觉关系的系统性理解。大量实验表明,Focus仅用30万条训练数据和20亿参数模型即达到当前最优性能,尤其在复杂界面中表现突出:在ScreenSpot上平均准确率达77.4%,在更具挑战性的ScreenSpot-Pro上提升13.3%。分析验证了双系统策略的有效性,展示了其在复杂界面交互中的应用潜力。
原文摘要 · Abstract (English)
Humans can flexibly switch between different modes of thinking based on task complexity: from rapid intuitive judgments to in-depth analytical understanding. However, current Graphical User Interface (GUI) grounding systems which locate interface elements based on natural language instructions rely solely on immediate prediction without reasoning, struggling to understand complex interface layouts with nested structures and hierarchical relationships, limiting their effectiveness on complex interfaces. Inspired by human dual-system cognition, we present Focus, a novel GUI grounding framework that combines fast prediction with systematic analysis. The framework dynamically switches between rapid and deliberate processing through an adaptive system switching based on task complexity, optimizing both efficiency and accuracy. Focus decomposes grounding into progressive stages: interface summarization, visual focused analysis, and precise coordinate prediction. This structured decomposition enables systematic understanding of both interface layouts and visual relationships. Extensive experiments show that Focus achieves state-of-the-art performance using only 300K of the training data with a 2B parameter model compared to existing approaches. Focus demonstrates superior performance particularly in complex GUI scenarios, achieving 77.4% average accuracy on ScreenSpot and 13.3% on the more challenging ScreenSpot-Pro. Our analysis reveals the effectiveness of this dual-system approach while demonstrating its potential for improving complex GUI interaction scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。