让电脑直接看屏幕操作界面,实现跨平台自动任务执行。
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

- 直接用图像处理界面,不依赖文字描述
- 在多个平台测试中表现超越现有方法
- 适合做自动化测试、辅助残障人士的工具
自动化图形用户界面(GUI)任务仍面临依赖文本表示、平台特异性操作空间和推理能力有限等挑战。我们提出Aguvis,一种统一的纯视觉框架,使自主GUI代理直接基于屏幕图像进行操作,标准化跨平台交互,并通过内部独白机制实现结构化推理。为此,我们构建了Aguvis数据集收集,包含多模态对齐与推理标注的大规模数据集,并设计两阶段训练流程,将GUI定位与规划推理分离。实验表明,Aguvis在离线与真实在线基准上均达到最先进水平,是首个完全自主、无需闭源模型的纯视觉GUI代理。项目开源所有数据集、模型与训练方案,网址:https://aguvis-project.github.io,以推动后续研究。
原文摘要 · Abstract (English)
Automating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities. We introduce Aguvis, a unified vision-based framework for autonomous GUI agents that directly operates on screen images, standardizes cross-platform interactions and incorporates structured reasoning via inner monologue. To enable this, we construct Aguvis Data Collection, a large-scale dataset with multimodal grounding and reasoning annotations, and develop a two-stage training pipeline that separates GUI grounding from planning and reasoning. Experiments show that Aguvis achieves state-of-the-art performance across offline and real-world online benchmarks, marking the first fully autonomous vision-based GUI agent that operates without closed-source models. We open-source all datasets, models, and training recipes at https://aguvis-project.github.io to advance future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。