系统梳理电脑使用智能体的现状与挑战,助力构建更实用的自动化工具。
A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions
- 提出三维度统一分类法:领域、交互方式、智能体能力
- 分析87个智能体与33个数据集,发现通用性差等六大短板
- 建议真实场景测试与标准化评估,推动技术落地
电脑使用智能体(ACUs)是一类能通过自然语言指令在桌面、手机和网页平台执行复杂任务的新兴系统,可通过鼠标点击、触屏操作等底层动作控制软件。尽管进展迅速,但其尚未成熟适用于日常使用。本综述全面分析了当前技术发展、趋势与研究空白。提出涵盖三个维度的统一分类体系:(I) 领域视角,刻画智能体运行环境;(II) 交互视角,描述观察模态(如截图、HTML)与动作模态(如鼠标、键盘、代码执行);(III) 智能体视角,解析其感知、推理与学习机制。基于该框架,评估了87个ACU与33个数据集,涵盖基于大模型与传统方法。识别出六大关键差距:泛化能力不足、学习效率低、规划能力弱、基准任务复杂度有限、评估标准不统一、研究与实际部署脱节。为此提出六项改进方向:(a) 使用视觉观测与底层控制提升泛化;(b) 实现超越静态提示的自适应学习;(c) 发展高效规划与推理方法;(d) 构建反映真实任务复杂度的基准;(e) 基于任务成功率建立标准化评估;(f) 将智能体设计与真实部署约束对齐。整体为迈向通用、鲁棒、可扩展的电脑使用智能体奠定基础。
原文摘要 · Abstract (English)
Agents for computer use (ACUs) are an emerging class of systems capable of executing complex tasks on digital devices -- such as desktops, mobile phones, and web platforms -- given instructions in natural language. These agents can automate tasks by controlling software via low-level actions like mouse clicks and touchscreen gestures. However, despite rapid progress, ACUs are not yet mature for everyday use. In this survey, we investigate the state-of-the-art, trends, and research gaps in the development of practical ACUs. We provide a comprehensive review of the ACU landscape, introducing a unifying taxonomy spanning three dimensions: (I) the domain perspective, characterizing agent operating contexts; (II) the interaction perspective, describing observation modalities (e.g., screenshots, HTML) and action modalities (e.g., mouse, keyboard, code execution); and (III) the agent perspective, detailing how agents perceive, reason, and learn. We review 87 ACUs and 33 datasets across foundation model-based and classical approaches through this taxonomy. Our analysis identifies six major research gaps: insufficient generalization, inefficient learning, limited planning, low task complexity in benchmarks, non-standardized evaluation, and a disconnect between research and practical conditions. To address these gaps, we advocate for: (a) vision-based observations and low-level control to enhance generalization; (b) adaptive learning beyond static prompting; (c) effective planning and reasoning methods and models; (d) benchmarks that reflect real-world task complexity; (e) standardized evaluation based on task success; (f) aligning agent design with real-world deployment constraints. Together, our taxonomy and analysis establish a foundation for advancing ACU research toward general-purpose agents for robust and scalable computer use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。