arXiv:2505.12370cs.AI2025-05NeurIPS被引 74

用强化学习让界面智能体更准识别按钮,3千样本超大模型表现

Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning

  • 通过自进化强化微调,用少量高质量数据训练界面理解模型
  • 70亿参数模型在高分辨率界面中达47.3%准确率,超越720亿参数模型
  • 适合做自动化测试、智能助手的开发者参考

图形用户界面(GUI)智能体在跨平台理解与执行用户指令方面已取得显著进展,但在复杂高分辨率专业环境中精准定位界面元素仍具挑战。传统监督微调方法需大量多样数据且泛化能力弱。为此,本文提出基于强化学习的框架,包含三项核心策略:(1) 精选种子数据以保证训练样本质量;(2) 使用密集策略梯度,依据预测准确率提供连续反馈;(3) 自进化强化微调机制,通过注意力图迭代优化模型。仅使用3000个训练样本,70亿参数模型在三个基线数据集上达到同类规模模型最优性能,尤其在ScreenSpot-Pro数据集上取得47.3%准确率,领先于720亿参数的UI-TARS-72B模型达24.2个百分点。结果表明,强化学习方法在高分辨率复杂环境下的界面理解任务中极具有效性。

原文摘要 · Abstract (English)

Graphical User Interface (GUI) agents have made substantial strides in understanding and executing user instructions across diverse platforms. Yet, grounding these instructions to precise interface elements remains challenging, especially in complex, high-resolution, professional environments. Traditional supervised finetuning (SFT) methods often require large volumes of diverse data and exhibit weak generalization. To overcome these limitations, we introduce a reinforcement learning (RL) based framework that incorporates three core strategies: (1) seed data curation to ensure high quality training samples, (2) a dense policy gradient that provides continuous feedback based on prediction accuracy, and (3) a self evolutionary reinforcement finetuning mechanism that iteratively refines the model using attention maps. With only 3k training samples, our 7B-parameter model achieves state-of-the-art results among similarly sized models on three grounding benchmarks. Notably, it attains 47.3\% accuracy on the ScreenSpot-Pro dataset, outperforming much larger models, such as UI-TARS-72B, by a margin of 24.2\%. These findings underscore the effectiveness of RL-based approaches in enhancing GUI agent performance, particularly in high-resolution, complex environments.

GUI智能体强化学习视觉定位自进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。