arXiv:2412.13501cs.AIcs.HC2024-12ACL综述被引 112

GUI智能体用大模型自动操作界面,像人一样点击、输入、导航。

GUI Agents: A Survey

  • 基于大模型的GUI智能体可自主完成界面交互任务。
  • 提出统一框架,涵盖感知、推理、规划与执行能力。
  • 适合对自动化人机交互感兴趣的开发者和研究者。

图形用户界面(GUI)智能体借助大型基础模型,成为自动化人机交互的变革性方法。这些智能体能通过GUI自主与数字系统或软件应用交互,模拟人类在不同平台上的点击、键入和导航等视觉操作。为应对日益增长的兴趣和其根本重要性,本文提供全面综述,对相关基准测试、评估指标、架构及训练方法进行分类。我们提出一个统一框架,明确其感知、推理、规划和执行能力。此外,识别出关键开放挑战并探讨未来重要方向。本工作为从业者和研究人员理解当前进展、技术、基准及亟待解决的核心问题提供了基础。

原文摘要 · Abstract (English)

Graphical User Interface (GUI) agents, powered by Large Foundation Models, have emerged as a transformative approach to automating human-computer interaction. These agents autonomously interact with digital systems or software applications via GUIs, emulating human actions such as clicking, typing, and navigating visual elements across diverse platforms. Motivated by the growing interest and fundamental importance of GUI agents, we provide a comprehensive survey that categorizes their benchmarks, evaluation metrics, architectures, and training methods. We propose a unified framework that delineates their perception, reasoning, planning, and acting capabilities. Furthermore, we identify important open challenges and discuss key future directions. Finally, this work serves as a basis for practitioners and researchers to gain an intuitive understanding of current progress, techniques, benchmarks, and critical open problems that remain to be addressed.

GUI智能体人机交互大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。