用大模型让电脑界面能听懂人话,自动操作软件。
Large Language Model-Brained GUI Agents: A Survey
- 用大语言模型理解界面元素和自然语言指令
- 可执行多步骤操作,支持网页、手机和桌面自动化
- 适合想提升效率的开发者与普通用户
GUI长期是人机交互的核心,提供直观的视觉操作方式。大语言模型(尤其是多模态模型)的出现开启了GUI自动化的新时代,其在自然语言理解、代码生成和视觉处理方面表现出色,催生了能够理解复杂界面并根据自然语言指令自主执行任务的「大模型驱动的GUI代理」。这类代理实现了从命令式操作到对话式交互的范式转变,使用户可通过简单对话完成复杂多步任务,应用场景涵盖网页浏览、移动应用操作和桌面自动化,显著提升用户体验。本文系统综述该领域,涵盖历史演进、核心组件与先进技术,探讨现有框架、训练数据收集与利用、专用大动作模型开发,以及评估指标与基准测试方法。同时分析新兴应用,识别关键研究空白,提出未来发展方向。本综述旨在为研究者与实践者提供知识整合与技术指引,推动该领域突破挑战,释放全部潜力。
原文摘要 · Abstract (English)
GUIs have long been central to human-computer interaction, providing an intuitive and visually-driven way to access and interact with digital systems. The advent of LLMs, particularly multimodal models, has ushered in a new era of GUI automation. They have demonstrated exceptional capabilities in natural language understanding, code generation, and visual processing. This has paved the way for a new generation of LLM-brained GUI agents capable of interpreting complex GUI elements and autonomously executing actions based on natural language instructions. These agents represent a paradigm shift, enabling users to perform intricate, multi-step tasks through simple conversational commands. Their applications span across web navigation, mobile app interactions, and desktop automation, offering a transformative user experience that revolutionizes how individuals interact with software. This emerging field is rapidly advancing, with significant progress in both research and industry. To provide a structured understanding of this trend, this paper presents a comprehensive survey of LLM-brained GUI agents, exploring their historical evolution, core components, and advanced techniques. We address research questions such as existing GUI agent frameworks, the collection and utilization of data for training specialized GUI agents, the development of large action models tailored for GUI tasks, and the evaluation metrics and benchmarks necessary to assess their effectiveness. Additionally, we examine emerging applications powered by these agents. Through a detailed analysis, this survey identifies key research gaps and outlines a roadmap for future advancements in the field. By consolidating foundational knowledge and state-of-the-art developments, this work aims to guide both researchers and practitioners in overcoming challenges and unlocking the full potential of LLM-brained GUI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。