arXiv:2602.16855cs.AIcs.CL2026-02被引 57

多平台通用图形界面智能体,支持跨设备实时协作与自动化操作。

Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents

  • 构建混合数据飞轮,结合仿真与云端沙盒提升训练数据质量。
  • 在20+基准上达成领先性能,移动端任务准确率达71.6。
  • 开源模型支持桌面、移动、浏览器等多平台,适合开发者部署使用。

本文提出GUI-Owl-1.5,最新一代原生图形界面智能体模型,具备多种尺寸(2B/4B/8B/32B/235B)的指令/思考变体,支持桌面、移动、浏览器等多种平台,实现云边协同与实时交互。该模型在超过20个开源基准上达到顶尖表现:在GUI自动化任务中,OSWorld得分为56.5,AndroidWorld达71.6,WebArena为48.4;在定位任务中,ScreenSpotPro得分为80.3;在工具调用任务中,OSWorld-MCP为47.6,MobileWorld为46.8;在记忆与知识任务中,GUI-Knowledge Bench得分为75.5。其关键技术包括:(1)混合数据飞轮:基于仿真环境与云端沙盒构建UI理解与轨迹生成的数据流水线,提升数据采集效率与质量;(2)统一能力增强:采用统一思维合成流程强化推理能力,重点优化工具调用、记忆与多智能体适应性;(3)多平台环境强化学习扩展:提出新型环境强化学习算法MRPO,解决多平台冲突与长序列任务训练效率低的问题。GUI-Owl-1.5已开源,线上云沙盒演示可在https://github.com/X-PLUG/MobileAgent获取。

原文摘要 · Abstract (English)

The paper introduces GUI-Owl-1.5, the latest native GUI agent model that features instruct/thinking variants in multiple sizes (2B/4B/8B/32B/235B) and supports a range of platforms (desktop, mobile, browser, and more) to enable cloud-edge collaboration and real-time interaction. GUI-Owl-1.5 achieves state-of-the-art results on more than 20+ GUI benchmarks on open-source models: (1) on GUI automation tasks, it obtains 56.5 on OSWorld, 71.6 on AndroidWorld, and 48.4 on WebArena; (2) on grounding tasks, it obtains 80.3 on ScreenSpotPro; (3) on tool-calling tasks, it obtains 47.6 on OSWorld-MCP, and 46.8 on MobileWorld; (4) on memory and knowledge tasks, it obtains 75.5 on GUI-Knowledge Bench. GUI-Owl-1.5 incorporates several key innovations: (1) Hybird Data Flywheel: we construct the data pipeline for UI understanding and trajectory generation based on a combination of simulated environments and cloud-based sandbox environments, in order to improve the efficiency and quality of data collection. (2) Unified Enhancement of Agent Capabilities: we use a unified thought-synthesis pipeline to enhance the model's reasoning capabilities, while placing particular emphasis on improving key agent abilities, including Tool/MCP use, memory and multi-agent adaptation; (3) Multi-platform Environment RL Scaling: We propose a new environment RL algorithm, MRPO, to address the challenges of multi-platform conflicts and the low training efficiency of long-horizon tasks. The GUI-Owl-1.5 models are open-sourced, and an online cloud-sandbox demo is available at https://github.com/X-PLUG/MobileAgent.

GUI智能体多平台自动化开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。