arXiv:2507.22025cs.AIcs.CL2025-07被引 33

提升GUI智能体的精准操作能力,解决视觉干扰与奖励设计难题。

UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding

  • 用连续奖励和思维简化机制优化训练过程
  • 在ScreenSpot-Pro上实现23%的精度提升
  • 适合需要高精度界面操作的研究与开发人员

多模态大模型的兴起推动了图形用户界面(GUI)智能体的发展。然而,现有训练与推理方法仍面临推理设计困境、奖励效率低下及视觉噪声问题。为此,我们提出UI-AGILE,在训练与推理两阶段增强GUI智能体性能。训练方面,改进监督微调流程:1)采用连续奖励函数激励高精度定位;2)引入“简单思考”奖励,平衡规划效率与定位准确率;3)基于裁剪的重采样策略缓解稀疏奖励问题,提升复杂任务学习效果。推理方面,提出分解式定位选择机制,将高分辨率图像分块处理,显著提升定位精度。实验表明,UI-AGILE在ScreenSpot-Pro与ScreenSpot-v2两个基准上达到当前最优定位表现,结合训练与推理优化方法,在ScreenSpot-Pro上较最佳基线提升23%定位准确率。代码已开源:https://github.com/KDEGroup/UI-AGILE。

原文摘要 · Abstract (English)

The emergence of Multimodal Large Language Models (MLLMs) has driven significant advances in Graphical User Interface (GUI) agent capabilities. Nevertheless, existing GUI agent training and inference techniques still suffer from a dilemma for reasoning designs, ineffective reward, and visual noise. To address these issues, we introduce UI-AGILE for enhancing GUI agents at both training and inference. For training, we propose a suite of improvements to the Supervised Fine-Tuning (SFT) process: 1) a continuous reward function to incentivize high-precision grounding; 2) a ``Simple Thinking'' reward to balance planning with speed and grounding accuracy; and 3) a cropping-based resampling strategy to mitigate the sparse reward problem and improve learning on complex tasks. For inference, we present decomposed grounding with selection to dramatically improve grounding accuracy on high-resolution displays by breaking the image into smaller, manageable parts. Experiments show that UI-AGILE achieves the state-of-the-art grounding performance on two benchmarks ScreenSpot-Pro and ScreenSpot-v2 while it also exhibits strong general agent capabilities. For instance, using both our training and inference enhancement methods brings 23\% grounding accuracy improvement over the best baseline on ScreenSpot-Pro. We provide the code in https://github.com/KDEGroup/UI-AGILE.

GUI智能体强化学习视觉定位大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。