用强化学习提升大模型在图形界面中的智能操作能力
A Survey on GUI Agents with Foundation Models Enhanced by Reinforcement Learning
- 将界面操作建模为马尔可夫决策过程,分感知-规划-执行模块构建框架
- 强化学习使代理在复杂环境中泛化能力显著提升,优于纯提示或微调方法
- 适合对自动化交互、AI助手研发感兴趣的开发者与研究者
图形用户界面(GUI)代理在多模态大语言模型(MLLM)驱动下,已成为实现与数字系统智能交互的有前景范式。本文系统综述了基于强化学习(RL)增强的GUI代理最新进展。首先将GUI代理任务形式化为马尔可夫决策过程,讨论典型执行环境与评估指标。随后回顾(M)LLM-based GUI代理的模块化架构,涵盖感知、规划和执行模块,并梳理代表性工作的演进路径。进一步将训练方法分为基于提示、监督微调(SFT)和强化学习三类,突出从简单提示工程到通过强化学习实现动态策略学习的演进。总结表明,多模态感知、决策推理与自适应动作生成的创新显著提升了代理在复杂真实环境中的泛化性与鲁棒性。最后指出关键挑战与未来方向,以构建更强大可靠的GUI代理。
原文摘要 · Abstract (English)
Graphical User Interface (GUI) agents, driven by Multi-modal Large Language Models (MLLMs), have emerged as a promising paradigm for enabling intelligent interaction with digital systems. This paper provides a structured survey of recent advances in GUI agents, focusing on architectures enhanced by Reinforcement Learning (RL). We first formalize GUI agent tasks as Markov Decision Processes and discuss typical execution environments and evaluation metrics. We then review the modular architecture of (M)LLM-based GUI agents, covering Perception, Planning, and Acting modules, and trace their evolution through representative works. Furthermore, we categorize GUI agent training methodologies into Prompt-based, Supervised Fine-Tuning (SFT)-based, and RL-based approaches, highlighting the progression from simple prompt engineering to dynamic policy learning via RL. Our summary illustrates how recent innovations in multimodal perception, decision reasoning, and adaptive action generation have significantly improved the generalization and robustness of GUI agents in complex real-world environments. We conclude by identifying key challenges and future directions for building more capable and reliable GUI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。