用大模型让软件操作自动化的智能代理,可像人一样点按钮、输文字。
GUI Agents with Foundation Models: A Comprehensive Survey
- 基于大模型理解界面,实现自动点击、输入等操作
- 梳理了关键数据集、统一框架与应用案例
- 适合研究AI自动化、人机交互的开发者和研究人员
近年来,基础模型(尤其是大语言模型LLM和多模态大语言模型MLLM)的发展推动了能执行复杂任务的智能代理诞生。通过利用(M)LLM处理和解析图形用户界面(GUI)的能力,这些代理可自主执行用户指令,模拟人类点击、输入等操作。本综述整合了基于(M)LLM的GUI代理的最新研究,重点突出数据资源、框架设计与应用创新。首先回顾代表性数据集与评测基准,随后提出一个通用统一框架,归纳先前研究的核心组件,并构建详细分类体系。此外,还探讨了相关商业应用。基于现有工作,我们识别出关键挑战并提出未来研究方向。希望本综述能推动该领域进一步发展。
原文摘要 · Abstract (English)
Recent advances in foundation models, particularly Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs), have facilitated the development of intelligent agents capable of performing complex tasks. By leveraging the ability of (M)LLMs to process and interpret Graphical User Interfaces (GUIs), these agents can autonomously execute user instructions, simulating human-like interactions such as clicking and typing. This survey consolidates recent research on (M)LLM-based GUI agents, highlighting key innovations in data resources, frameworks, and applications. We begin by reviewing representative datasets and benchmarks, followed by an overview of a generalized, unified framework that encapsulates the essential components of prior studies, supported by a detailed taxonomy. Additionally, we explore relevant commercial applications. Drawing insights from existing work, we identify key challenges and propose future research directions. We hope this survey will inspire further advancements in the field of (M)LLM-based GUI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。