arXiv:2508.03700cs.HCcs.AI2025-08被引 16

MagicGUI让手机界面智能体更懂操作,能看会想还能自学习。

MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning

  • 用大规模数据管道构建多模态界面数据集,支持精准识别与定位
  • 通过强化学习训练,在多个基准上实现领先表现,泛化能力强
  • 适合开发移动智能助手、自动化测试工具的研发人员参考

本文提出MagicGUI,一个面向真实移动界面环境的通用型智能体框架,解决感知、定位与推理难题。其核心由六大组件构成:(1)通过可扩展的GUI数据管道,聚合开源仓库、自动爬取和人工标注,构建迄今最大最多样化的以界面为中心的多模态数据集;(2)增强的感知与定位能力,实现细粒度的多模态对齐,精准理解界面元素;(3)统一的动作空间,涵盖基础操作与复杂交互意图,支持人机协作;(4)面向规划的推理机制,可将复杂指令分解为带中间元计划的序列动作;(5)两阶段迭代训练流程,先在780万样本上持续预训练,再使用空间增强复合奖励与双过滤策略进行强化微调;(6)在自有Magic-RICH基准及十余个公开基准上表现优异,展现强大泛化能力与实际部署潜力。

原文摘要 · Abstract (English)

This paper presents MagicGUI, a foundational mobile GUI agent designed to address critical challenges in perception, grounding, and reasoning within real-world mobile GUI environments. The framework is underpinned by following six key components: (1) a comprehensive and accurate dataset, constructed via the scalable GUI Data Pipeline, which aggregates the largest and most diverse GUI-centric multimodal data to date from open-source repositories, automated crawling, and targeted manual annotation; (2) enhanced perception and grounding capabilities, facilitating fine-grained multimodal alignment for UI element referencing, grounding, and screen comprehension; (3) a comprehensive and unified action space, encompassing both fundamental UI operations and complex interactive intents to support human-agent interactions; (4) planning-oriented reasoning mechanisms that enable the model to decompose complex user instructions into sequential actions with explicit intermediate meta-paln reasoning; (5) an iterative two-stage training procedure, combining large-scale continue pre-training on 7.8M samples with reinforcement fine-tuning utilizing a spatially enhanced composite reward and dual filtering strategy; and (6) competitive performance on both the proprietary Magic-RICH benchmark and over a dozen public benchmarks, achieving superior performance across GUI perception and agent tasks, while demonstrating robust generalization and real-world deployment potential in practical mobile GUI scenarios, as detailed in Figure 1.

移动智能体界面理解强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。