让手机界面智能体更高效:数据与环境协同扩增,提升长程决策能力。
HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents

- 构建数据-环境双轮驱动的扩展框架,解决移动端交互瓶颈。
- 在2000+设备上部署百万级动作数据,任务覆盖超3.4万项。
- 适合研究移动机器人、自动化测试及长程决策系统的人群。
随着大模型从内容理解转向数字环境操作,移动端图形用户界面(GUI)成为数字具身智能的重要挑战场景。移动智能体面临三大耦合约束:对复杂界面的精准感知、高质量交互数据的可扩展获取、以及累积执行错误下的鲁棒长程决策。本文提出HyMobileAgent,基于Hy3.0-VL-A3B这一视觉原生基础模型,具备任意分辨率输入、A3B规模部署预算和32K上下文窗口,以建模长期交互历史。不依赖单纯模型扩容,而是设计数据与环境协同扩展框架:集成基于模拟界面生成、拒绝采样和图标特异性增强的界面感知飞轮;将教程视频转化为结构化交互数据的知识管道;在2000多个沙盒与真实设备实例上部署的百万级动作数据流水线,并实现自动失败归因;提供34个模拟应用、超过34000项任务的可重置训练环境PhoneWorld Mock App Factory;以及带显式死循环检测的结构化规划与反思机制,保障长程执行可靠性。此外,引入渐进式训练方案,包括中期监督微调与任务特异奖励设计的强化学习。
原文摘要 · Abstract (English)
As large multimodal models move from understanding content to operating on digital environments, mobile GUI has emerged as a challenging and consequential testbed for digital embodied intelligence. Mobile agents operate under three coupled constraints: precise perception of complex interfaces, scalable acquisition of high-quality interaction data, and robust long-horizon decision making under compounding execution errors. This report presents HyMobileAgent, a mobile GUI agent built on Hy3.0-VL-A3B, a vision-native foundation model featuring native any-resolution input, an A3B-scale deployment budget, and a 32K context window to model extended interaction histories. Rather than relying solely on model scaling, we develop a joint data and environment centric scaling framework to address the key bottlenecks of mobile interaction. Our framework integrates a GUI perception flywheel combining mock-interface synthesis, rejection sampling, and icon-specific augmentation; a knowledge pipeline that transforms tutorial videos into structured interaction data; a million-scale action data pipeline deployed across more than 2000 sandbox and real-device instances with automated failure attribution; the PhoneWorld Mock App Factory, providing a resettable training environment with 34 mock applications and over 34000 tasks; and a structured Planning-and-Reflection mechanism with explicit dead-loop detection for reliable long-horizon execution. We also introduce a progressive training recipe consisting of mid-training, supervised fine-tuning, and reinforcement learning with task-specific reward designs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。