arXiv:2508.20018cs.AIcs.CL2025-08被引 1

提出分阶段交错强化学习框架,提升移动端GUI多智能体协作效率

SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control

  • 分步更新单个智能体,保持其他固定,实现稳定训练
  • 在移动端GUI任务中显著优于现有方法,高阶与低阶任务均表现优异
  • 适用于多智能体数学推理等复杂场景,具通用性

大型视觉语言模型(LVLM)和智能体系统的快速发展,激发了可将自然语言可靠转化为界面操作的移动GUI智能体研究兴趣。然而,现有单智能体方法受结构限制;尽管多智能体系统能天然解耦不同能力,但当前多智能体强化学习(MARL)常因效率低下且不兼容现有LVLM架构而受限。为此,本文提出SWIRL——一种面向多智能体系统的分阶段交错强化学习工作流。SWIRL将MARL重构为一系列单智能体强化学习任务,每次仅更新一个智能体,其余保持固定,从而实现稳定训练并促进高效协同。理论上,本文提供分步安全边界、跨轮单调改进定理及回报收敛保证,确保优化过程稳健且有理论依据。应用于移动端GUI控制时,SWIRL构建了导航器(Navigator),将语言与屏幕上下文转化为结构化计划;以及执行器(Interactor),将计划映射为可执行原子动作。大量实验表明,该方法在高阶与低阶GUI基准测试中均表现卓越。此外,SWIRL在多智能体数学推理任务中也展现出强大能力,凸显其作为高效、鲁棒多智能体系统通用框架的潜力。

原文摘要 · Abstract (English)

The rapid advancement of large vision language models (LVLMs) and agent systems has heightened interest in mobile GUI agents that can reliably translate natural language into interface operations. Existing single-agent approaches, however, remain limited by structural constraints. Although multi-agent systems naturally decouple different competencies, recent progress in multi-agent reinforcement learning (MARL) has often been hindered by inefficiency and remains incompatible with current LVLM architectures. To address these challenges, we introduce SWIRL, a staged workflow for interleaved reinforcement learning designed for multi-agent systems. SWIRL reformulates MARL into a sequence of single-agent reinforcement learning tasks, updating one agent at a time while keeping the others fixed. This formulation enables stable training and promotes efficient coordination across agents. Theoretically, we provide a stepwise safety bound, a cross-round monotonic improvement theorem, and convergence guarantees on return, ensuring robust and principled optimization. In application to mobile GUI control, SWIRL instantiates a Navigator that converts language and screen context into structured plans, and an Interactor that grounds these plans into executable atomic actions. Extensive experiments demonstrate superior performance on both high-level and low-level GUI benchmarks. Beyond GUI tasks, SWIRL also demonstrates strong capability in multi-agent mathematical reasoning, underscoring its potential as a general framework for developing efficient and robust multi-agent systems.

多智能体强化学习GUI控制LVLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。