arXiv:2508.04482cs.AIcs.CL2025-08综述被引 54

综述基于大模型的系统级智能代理,解析其构建与评估方法。

OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use

  • 以操作系统界面为交互环境,构建可执行任务的AI代理
  • 提出代理核心组件框架,涵盖观察、行动与任务规划能力
  • 适合研究通用AI助手的学者及工业界开发者参考

实现如《钢铁侠》中贾维斯般全能智能助手的梦想正因多模态大语言模型(M)LLMs的发展而加速。基于(M)LLMs的系统级代理(OS Agents)通过操作计算机或手机等通用设备的图形用户界面(GUI),在操作系统提供的环境中自动化任务,已取得显著进展。本文全面综述此类先进代理,首先阐明其基础概念,包括环境、观测空间与动作空间,并梳理理解、规划与具身化等关键能力。接着分析构建方法,聚焦领域专用基础模型与代理框架;深入评述评估协议与基准测试,展示代理在多样化任务中的表现。最后讨论当前挑战与未来方向,如安全隐私、个性化与自我进化。本综述旨在整合现有研究成果,推动学术与产业协同发展。相关开源项目持续更新,9页版本已获ACL 2025接收,提供领域概览。

原文摘要 · Abstract (English)

The dream to create AI assistants as capable and versatile as the fictional J.A.R.V.I.S from Iron Man has long captivated imaginations. With the evolution of (multi-modal) large language models ((M)LLMs), this dream is closer to reality, as (M)LLM-based Agents using computing devices (e.g., computers and mobile phones) by operating within the environments and interfaces (e.g., Graphical User Interface (GUI)) provided by operating systems (OS) to automate tasks have significantly advanced. This paper presents a comprehensive survey of these advanced agents, designated as OS Agents. We begin by elucidating the fundamentals of OS Agents, exploring their key components including the environment, observation space, and action space, and outlining essential capabilities such as understanding, planning, and grounding. We then examine methodologies for constructing OS Agents, focusing on domain-specific foundation models and agent frameworks. A detailed review of evaluation protocols and benchmarks highlights how OS Agents are assessed across diverse tasks. Finally, we discuss current challenges and identify promising directions for future research, including safety and privacy, personalization and self-evolution. This survey aims to consolidate the state of OS Agents research, providing insights to guide both academic inquiry and industrial development. An open-source GitHub repository is maintained as a dynamic resource to foster further innovation in this field. We present a 9-page version of our work, accepted by ACL 2025, to provide a concise overview to the domain.

智能代理大模型系统自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。