arXiv:2508.15144cs.AI2025-08被引 170

开源通用GUI自动化框架,性能刷新多平台基准

Mobile-Agent-v3: Fundamental Agents for GUI Automation

  • 构建云端跨平台虚拟环境,自动生成高质量交互数据并迭代优化
  • 在安卓和桌面系统上分别达73.3和37.7分,超越现有开源模型
  • 支持模块化部署,适合研究与工业级自动化场景

本文提出GUI-Owl,一个基础性GUI代理模型,在涵盖桌面与移动平台的十个GUI基准测试中表现优异,覆盖视觉定位、问答、规划、决策与程序知识。GUI-Owl-7B在AndroidWorld上取得66.4分,在OSWorld上为29.4分。在此基础上,我们推出Mobile-Agent-v3通用GUI代理框架,性能进一步提升至AndroidWorld 73.3分,OSWorld 37.7分,创下开源框架新纪录。GUI-Owl包含三大创新:(1) 大规模环境基础设施:基于云的虚拟环境覆盖Android、Ubuntu、macOS与Windows,支持自动查询生成与正确性验证,实现自我演进的交互轨迹生成;(2) 多样化的基础代理能力:整合界面定位、规划、动作语义与推理模式,支持端到端决策,并可作为多智能体系统的模块组件;(3) 可扩展环境强化学习:提出全异步训练的强化学习框架与轨迹感知相对策略优化(TRPO),在线学习下在OSWorld达到34.9分。GUI-Owl与Mobile-Agent-v3已在GitHub开源:https://github.com/X-PLUG/MobileAgent。

原文摘要 · Abstract (English)

This paper introduces GUI-Owl, a foundational GUI agent model that achieves state-of-the-art performance among open-source end-to-end models on ten GUI benchmarks across desktop and mobile environments, covering grounding, question answering, planning, decision-making, and procedural knowledge. GUI-Owl-7B achieves 66.4 on AndroidWorld and 29.4 on OSWorld. Building on this, we propose Mobile-Agent-v3, a general-purpose GUI agent framework that further improves performance to 73.3 on AndroidWorld and 37.7 on OSWorld, setting a new state-of-the-art for open-source GUI agent frameworks. GUI-Owl incorporates three key innovations: (1) Large-scale Environment Infrastructure: a cloud-based virtual environment spanning Android, Ubuntu, macOS, and Windows, enabling our Self-Evolving GUI Trajectory Production framework. This generates high-quality interaction data via automated query generation and correctness validation, leveraging GUI-Owl to refine trajectories iteratively, forming a self-improving loop. It supports diverse data pipelines and reduces manual annotation. (2) Diverse Foundational Agent Capabilities: by integrating UI grounding, planning, action semantics, and reasoning patterns, GUI-Owl supports end-to-end decision-making and can act as a modular component in multi-agent systems. (3) Scalable Environment RL: we develop a scalable reinforcement learning framework with fully asynchronous training for real-world alignment. We also introduce Trajectory-aware Relative Policy Optimization (TRPO) for online RL, achieving 34.9 on OSWorld. GUI-Owl and Mobile-Agent-v3 are open-sourced at https://github.com/X-PLUG/MobileAgent.

GUI自动化多智能体强化学习开源框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。