arXiv:2504.13865cs.HCcs.AI2025-04综述被引 65

综述大模型驱动的图形界面智能体技术进展与挑战

A Survey on (M)LLM-Based GUI Agents

  • 拆解四大核心模块:感知、探索、规划、交互
  • 揭示多模态理解与长程规划是关键突破点
  • 适合研究人机交互与自动化系统的开发者参考

图形用户界面(GUI)智能体正成为人机交互的革新范式,从规则化脚本演进为能理解并执行复杂操作的AI系统。本文系统综述基于大语言模型(LLM)的GUI智能体,分析其架构基础、技术组件与评估方法。识别出四大核心构成:(1) 融合文本解析与多模态理解的感知系统;(2) 通过内部建模、历史经验与外部检索构建知识库的探索机制;(3) 借助高级推理进行任务分解与执行的规划框架;(4) 具备安全控制的能力的动作生成交互系统。研究表明,大模型与多模态学习的进展显著推动了桌面、移动及网页平台的智能自动化。文中批判性审视现有评估框架,指出基准测试的方法局限,并提出标准化方向。同时识别出准确元素定位、有效知识检索、长程规划与安全执行等关键技术挑战,展望未来研究路径。本综述为研究人员与实践者提供领域现状全景图与未来发展方向。

原文摘要 · Abstract (English)

Graphical User Interface (GUI) Agents have emerged as a transformative paradigm in human-computer interaction, evolving from rule-based automation scripts to sophisticated AI-driven systems capable of understanding and executing complex interface operations. This survey provides a comprehensive examination of the rapidly advancing field of LLM-based GUI Agents, systematically analyzing their architectural foundations, technical components, and evaluation methodologies. We identify and analyze four fundamental components that constitute modern GUI Agents: (1) perception systems that integrate text-based parsing with multimodal understanding for comprehensive interface comprehension; (2) exploration mechanisms that construct and maintain knowledge bases through internal modeling, historical experience, and external information retrieval; (3) planning frameworks that leverage advanced reasoning methodologies for task decomposition and execution; and (4) interaction systems that manage action generation with robust safety controls. Through rigorous analysis of these components, we reveal how recent advances in large language models and multimodal learning have revolutionized GUI automation across desktop, mobile, and web platforms. We critically examine current evaluation frameworks, highlighting methodological limitations in existing benchmarks while proposing directions for standardization. This survey also identifies key technical challenges, including accurate element localization, effective knowledge retrieval, long-horizon planning, and safety-aware execution control, while outlining promising research directions for enhancing GUI Agents' capabilities. Our systematic review provides researchers and practitioners with a thorough understanding of the field's current state and offers insights into future developments in intelligent interface automation.

GUI智能体大模型人机交互自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。