用强化学习让AI更懂图形界面操作,仅需少量数据就能搞定复杂任务。
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
- 基于统一动作空间规则建模,用强化学习提升视觉语言模型的界面操作能力。
- 仅用3000条数据(原方法1300万条的0.02%)就在8个基准上超越现有水平。
- 适合想在真实场景中部署高阶自动化工具的研究者与开发者。
当前构建图形用户界面(GUI)智能体主要依赖在大视觉语言模型(LVLM)上进行监督微调,但该方法需大量训练数据,且难以理解截图和泛化到未见界面,严重限制了其在真实场景中的应用,尤其在高阶任务上。受大推理模型中强化微调(RFT,如DeepSeek-R1)的启发,我们提出首个面向高阶现实任务的、基于统一动作空间规则建模的强化学习框架 ame,以增强LVLM的GUI能力。通过在多个平台(包括Windows、Linux、macOS、Android和Web)上使用少量精心筛选的高质量数据,并采用组相对策略优化(GRPO)等算法更新模型, ame 在跨移动、桌面、网页三类平台的八个基准测试中,仅用3000条数据(相较之前最佳方法OS-Atlas的1300万条,仅为0.02%),便实现了显著更优的表现。结果表明,基于统一动作空间规则建模的强化学习,在提升LVLM执行真实世界GUI任务的能力方面具有巨大潜力。
原文摘要 · Abstract (English)
Existing efforts in building Graphical User Interface (GUI) agents largely rely on the training paradigm of supervised fine-tuning on Large Vision-Language Models (LVLMs). However, this approach not only demands extensive amounts of training data but also struggles to effectively understand GUI screenshots and generalize to unseen interfaces. The issue significantly limits its application in real-world scenarios, especially for high-level tasks. Inspired by Reinforcement Fine-Tuning (RFT) in large reasoning models (e.g., DeepSeek-R1), which efficiently enhances the problem-solving capabilities of large language models in real-world settings, we propose \name, the first reinforcement learning framework designed to enhance the GUI capabilities of LVLMs in high-level real-world task scenarios, through unified action space rule modeling. By leveraging a small amount of carefully curated high-quality data across multiple platforms (including Windows, Linux, MacOS, Android, and Web) and employing policy optimization algorithms such as Group Relative Policy Optimization (GRPO) to update the model, \name achieves superior performance using only 0.02\% of the data (3K vs. 13M) compared to previous state-of-the-art methods like OS-Atlas across eight benchmarks spanning three different platforms (mobile, desktop, and web). These results demonstrate the immense potential of reinforcement learning based on unified action space rule modeling in improving the execution capabilities of LVLMs for real-world GUI agent tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。