让AI看懂界面时能自动聚焦重点区域,复杂任务分步分析
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
- 根据任务难度自动切分视觉区域,动态调整推理阶段
- 7B模型在ScreenSpot-Pro上达60.8%准确率,领先同类模型
- 适合需要精准界面定位的自动化测试与智能助手场景
现有GUI定位方法在高分辨率截图中常难以实现细粒度定位。为此,我们提出GUI-ARP框架,支持自适应多阶段推理。通过提出的自适应区域感知(ARP)和自适应阶段控制(ASC),GUI-ARP动态利用视觉注意力裁剪任务相关区域,并自适应调整推理策略:简单任务采用单阶段推理,复杂场景则进行多阶段分析。该方法基于两阶段训练流程,融合监督微调与基于组相对策略优化(GRPO)的强化学习微调。大量实验表明,GUI-ARP在挑战性GUI定位基准上达到顶尖性能,7B模型在ScreenSpot-Pro上达到60.8%准确率,在UI-Vision上达30.9%。值得注意的是,GUI-ARP-7B在性能上已可媲美开源72B模型(UI-TARS-72B为38.1%)及专有模型。
原文摘要 · Abstract (English)
Existing GUI grounding methods often struggle with fine-grained localization in high-resolution screenshots. To address this, we propose GUI-ARP, a novel framework that enables adaptive multi-stage inference. Equipped with the proposed Adaptive Region Perception (ARP) and Adaptive Stage Controlling (ASC), GUI-ARP dynamically exploits visual attention for cropping task-relevant regions and adapts its inference strategy, performing a single-stage inference for simple cases and a multi-stage analysis for more complex scenarios. This is achieved through a two-phase training pipeline that integrates supervised fine-tuning with reinforcement fine-tuning based on Group Relative Policy Optimization (GRPO). Extensive experiments demonstrate that the proposed GUI-ARP achieves state-of-the-art performance on challenging GUI grounding benchmarks, with a 7B model reaching 60.8% accuracy on ScreenSpot-Pro and 30.9% on UI-Vision benchmark. Notably, GUI-ARP-7B demonstrates strong competitiveness against open-source 72B models (UI-TARS-72B at 38.1%) and proprietary models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。