arXiv:2512.02423cs.CV2025-12NeurIPS被引 4

用多轮强化学习提升智能体在复杂界面中的导航能力。

GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning

  • 构建可自定义的GUI仿真环境,支持灵活配置界面元素和导航图。
  • 多轮强化学习显著提升智能体在未知场景下的探索与导航性能。
  • 适合研究通用界面智能体或需自主探索的应用场景。

随着大视觉语言模型的发展,图形用户界面(GUI)智能体任务从单屏操作转向复杂的屏幕导航挑战。然而,现实世界中的GUI环境(如桌面软件和移动App)通常复杂且专有,难以获取完整的环境信息,限制了智能体训练与评估的系统性研究。为此,我们提出GUI Exploration Lab,一个支持灵活定义屏幕、图标与导航图的仿真环境引擎,并提供完整的环境信息访问,以实现全面的智能体训练与评估。大量实验表明,监督微调可有效记忆基础知识,为后续训练奠定基础;单轮强化学习提升了对未见场景的泛化能力;而多轮强化学习通过交互式试错促进探索策略发展,进一步优化了屏幕导航表现。我们在静态与动态基准上验证方法,结果表明其能有效推广至真实场景。这些发现凸显了强化学习在GUI导航中的优势,为构建更强大、通用的智能体提供了实用指导。

原文摘要 · Abstract (English)

With the rapid development of Large Vision Language Models, the focus of Graphical User Interface (GUI) agent tasks shifts from single-screen tasks to complex screen navigation challenges. However, real-world GUI environments, such as PC software and mobile Apps, are often complex and proprietary, making it difficult to obtain the comprehensive environment information needed for agent training and evaluation. This limitation hinders systematic investigation and benchmarking of agent navigation capabilities. To address this limitation, we introduce GUI Exploration Lab, a simulation environment engine for GUI agent navigation research that enables flexible definition and composition of screens, icons, and navigation graphs, while providing full access to environment information for comprehensive agent training and evaluation. Through extensive experiments, we find that supervised fine-tuning enables effective memorization of fundamental knowledge, serving as a crucial foundation for subsequent training. Building on this, single-turn reinforcement learning further enhances generalization to unseen scenarios. Finally, multi-turn reinforcement learning encourages the development of exploration strategies through interactive trial and error, leading to further improvements in screen navigation performance. We validate our methods on both static and interactive benchmarks, demonstrating that our findings generalize effectively to real-world scenarios. These findings demonstrate the advantages of reinforcement learning approaches in GUI navigation and offer practical guidance for building more capable and generalizable GUI agents.

GUI导航强化学习智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。