MAI-UI打造真实场景通用GUI智能体,提升交互能力与部署效率。
MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
- 构建自进化数据流与设备云协同系统,支持复杂用户交互。
- 在多平台导航任务中达76.7%成功率,刷新移动GUI领域新纪录。
- 兼顾隐私保护与低延迟,适合实际落地的智能助手应用。
GUI智能体有望重塑人机交互。我们提出MAI-UI系列基础GUI代理,涵盖2B、8B、32B及235B-A22B等规模模型。针对真实部署中的四大挑战——缺乏自然人机交互、仅依赖界面操作、缺少实用部署架构、动态环境鲁棒性差——我们设计统一方案:自演化数据管道扩展导航数据至包含用户行为与MCP工具调用;原生设备-云协作系统按任务状态调度执行;在线强化学习框架通过并行环境扩增与上下文长度优化实现高效训练。在基准测试中,其在ScreenSpot-Pro上达73.5%,优于Gemini-3-Pro和Seed1.8;MMBench GUI L2达91.3%,OSWorld-G为70.9%,UI-Vision为49.2%。移动导航任务中,AndroidWorld成功率达76.7%,超越UI-Tars-2、Gemini-2.5-Pro和Seed1.8;MobileWorld达41.7%,显著优于端到端模型且接近Gemini-3-Pro框架。在线强化学习实验显示,将并行环境从32增至512可提升5.2分,步数预算从15增至50提升4.3分。设备云协作系统使本地性能提升33%,云调用减少超40%,保障用户隐私。
原文摘要 · Abstract (English)
The development of GUI agents could revolutionize the next generation of human-computer interaction. Motivated by this vision, we present MAI-UI, a family of foundation GUI agents spanning the full spectrum of sizes, including 2B, 8B, 32B, and 235B-A22B variants. We identify four key challenges to realistic deployment: the lack of native agent-user interaction, the limits of UI-only operation, the absence of a practical deployment architecture, and brittleness in dynamic environments. MAI-UI addresses these issues with a unified methodology: a self-evolving data pipeline that expands the navigation data to include user interaction and MCP tool calls, a native device-cloud collaboration system routes execution by task state, and an online RL framework with advanced optimizations to scale parallel environments and context length. MAI-UI establishes new state-of-the-art across GUI grounding and mobile navigation. On grounding benchmarks, it reaches 73.5% on ScreenSpot-Pro, 91.3% on MMBench GUI L2, 70.9% on OSWorld-G, and 49.2% on UI-Vision, surpassing Gemini-3-Pro and Seed1.8 on ScreenSpot-Pro. On mobile GUI navigation, it sets a new SOTA of 76.7% on AndroidWorld, surpassing UI-Tars-2, Gemini-2.5-Pro and Seed1.8. On MobileWorld, MAI-UI obtains 41.7% success rate, significantly outperforming end-to-end GUI models and competitive with Gemini-3-Pro based agentic frameworks. Our online RL experiments show significant gains from scaling parallel environments from 32 to 512 (+5.2 points) and increasing environment step budget from 15 to 50 (+4.3 points). Finally, the native device-cloud collaboration system improves on-device performance by 33%, reduces cloud model calls by over 40%, and preserves user privacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。