让轻量级GUI代理通过多角色协作实现高效自动化
Towards Scalable Lightweight GUI Agents via Multi-role Orchestration

- 用角色导向数据合成与两阶段训练提升轻量模型能力
- LAMO-3B支持单体与多智能体协同,任务扩展性显著增强
- 适合资源受限设备上需要灵活扩展的GUI自动化场景
由多模态大语言模型驱动的自主图形用户界面(GUI)代理可实现终端设备上的数字自动化。尽管参数与数据规模扩大带来了显著提升,先进方法在资源受限设备上仍面临高昂部署成本。面对复杂的现实场景,轻量级GUI代理受限于能力不足与端到端持续学习下的任务可扩展性差,难以适应多智能体系统(MAS),且训练多个专用技能专家成本过高。能否在成本与可扩展性间取得有效平衡,使轻量级多模态大模型参与真实GUI工作流?为此,我们提出LAMO框架,赋予轻量级多模态大模型特定于GUI的知识与任务可扩展性,通过多角色编排扩展其能力边界。LAMO结合角色导向数据合成与两阶段训练:(i) 采用困惑度加权交叉熵优化进行监督微调,实现知识蒸馏与视觉感知增强;(ii) 通过强化学习实现角色导向的协同探索。基于LAMO,我们构建了任务可扩展的原生GUI代理LAMO-3B,支持单体执行与MAS式编排。当与先进规划器配合作为即插即用策略执行器时,LAMO-3B能持续受益于规划器进步,突破性能上限。静态与在线评估均验证了设计的有效性。
原文摘要 · Abstract (English)
Autonomous Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) enable digital automation on end-user devices. While scaling both parameters and data has yielded substantial gains, advanced methods still suffer from prohibitive deployment costs on resource-constrained devices. When facing complex in-the-wild scenarios, lightweight GUI agents are bottlenecked by limited capacity and poor task scalability under end-to-end episodic learning, impeding adaptation to multi-agent systems (MAS), while training multiple skill-specific experts remains costly. Can we strike an effective trade-off in this cost-scalability dilemma, enabling lightweight MLLMs to participate in realistic GUI workflows? To address these challenges, we propose the LAMO framework, which endows a lightweight MLLM with GUI-specific knowledge and task scalability, allowing multi-role orchestration to expand its capability boundary for GUI automation. LAMO combines role-oriented data synthesis with a two-stage training recipe: (i) supervised fine-tuning with Perplexity-Weighted Cross-Entropy optimization for knowledge distillation and visual perception enhancement, and (ii) reinforcement learning for role-oriented cooperative exploration. With LAMO, we develop a task-scalable native GUI agent, LAMO-3B, supporting monolithic execution and MAS-style orchestration. When paired with advanced planners as a plug-and-play policy executor, LAMO-3B can continuously benefit from planner advances, enabling a higher performance ceiling. Extensive static and online evaluations validate the effectiveness of our design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。