arXiv:2507.05720cs.LGcs.CL2025-07被引 45

让手机界面智能体在线学习,提升泛化能力与稳定性。

MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment

  • 通过自探索生成可学习的任务课程,实现在线训练
  • 在三个基准上显著提升任务成功率与执行效率
  • 适合需要适应复杂多变手机界面的自动化研究者

近年来,基于视觉的GUI智能体被广泛用于自动化日常移动和网页任务。这些智能体通过解析原始屏幕截图,自主决定点击、滚动或输入位置,无需依赖手工规则或应用特定API。然而,大多数现有方法在离线环境中使用预收集轨迹训练,导致扩展性差、对特定界面模板过拟合,且在面对未见环境时策略脆弱。我们提出MobileGUI-RL,一种在在线环境中训练GUI智能体的可扩展框架。该框架包含两个关键组件:(i) 通过自探索与筛选生成可学习的任务课程;(ii) 将GRPO改进为适用于GUI导航,采用轨迹感知优势和复合奖励,平衡任务成功与执行效率。在三个在线移动智能体基准上的实验显示一致性能提升,验证了该方法的有效性。

原文摘要 · Abstract (English)

Recently, there has been a surge of vision-based GUI agents designed to automate everyday mobile and web tasks. These agents interpret raw GUI screenshots and autonomously decide where to click, scroll, or type, which bypasses handcrafted rules and app-specific APIs. However, most existing methods trained GUI agent in the offline environment using pre-collected trajectories. This approach limits scalability, causes overfitting to specific UI templates, and leads to brittle policies when faced with unseen environment. We present MobileGUI-RL, a scalable framework that trains GUI agent in online environment. MobileGUI-RL contains two key components. It (i) synthesizes a curriculum of learnable tasks through self-exploration and filtering, and (ii) adapts GRPO to GUI navigation with trajectory-aware advantages and composite rewards that balance task success and execution efficiency. Experiments on three online mobile-agent benchmarks show consistent gains, validating the effectiveness of our approach.

GUI智能体强化学习在线训练移动端自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。