构建首个覆盖全屏操作的桌面智能体评测基准,推动自动化办公研究
GUI-360$^\circ$: A Comprehensive Dataset and Benchmark for Computer-Using Agents
- 用大模型自动构建真实办公场景任务与多模态轨迹数据
- 含超120万步操作记录,支持图形界面定位、屏幕解析和动作预测
- 适合研究桌面智能体、人机交互及自动化工具的开发者使用
我们提出GUI-360°,一个大规模综合性数据集与评测基准,旨在推动计算机使用智能体(CUAs)的发展。现有研究面临三大瓶颈:真实世界CUA任务稀缺、多模态轨迹自动化采集与标注管道缺失、缺乏统一评估GUI定位、屏幕解析与动作预测的基准。GUI-360°通过大模型增强的自动化流程,实现查询获取、环境模板构建、任务实例化、批量执行及大模型驱动的质量过滤。数据集包含超过120万步已执行操作,涵盖数千条轨迹,覆盖主流Windows办公应用,提供全分辨率截图、可访问性元数据、目标任务、中间推理轨迹,以及成功与失败的操作序列。支持三种典型任务:GUI定位、屏幕解析、动作预测,以及融合GUI与API的混合动作空间。在该数据集上对先进视觉-语言模型的评测显示,其在定位与动作预测方面存在显著不足;监督微调与强化学习虽带来提升,但仍未达到人类水平可靠性。我们已公开发布GUI-360°及配套代码,促进可复现研究,加速鲁棒桌面智能体的发展。完整数据集已上线Hugging Face:https://huggingface.co/datasets/vyokky/GUI-360。
原文摘要 · Abstract (English)
We introduce GUI-360$^\circ$, a large-scale, comprehensive dataset and benchmark suite designed to advance computer-using agents (CUAs). CUAs present unique challenges and is constrained by three persistent gaps: a scarcity of real-world CUA tasks, the lack of automated collection-and-annotation pipelines for multi-modal trajectories, and the absence of a unified benchmark that jointly evaluates GUI grounding, screen parsing, and action prediction. GUI-360$^\circ$ addresses these gaps with an LLM-augmented, largely automated pipeline for query sourcing, environment-template construction, task instantiation, batched execution, and LLM-driven quality filtering. The released corpus contains over 1.2M executed action steps across thousands of trajectories in popular Windows office applications, and includes full-resolution screenshots, accessibility metadata when available, instantiated goals, intermediate reasoning traces, and both successful and failed action trajectories. The dataset supports three canonical tasks, GUI grounding, screen parsing, and action prediction, and a hybrid GUI+API action space that reflects modern agent designs. Benchmarking state-of-the-art vision--language models on GUI-360$^\circ$ reveals substantial out-of-the-box shortcomings in grounding and action prediction; supervised fine-tuning and reinforcement learning yield significant gains but do not close the gap to human-level reliability. We release GUI-360$^\circ$ and accompanying code to facilitate reproducible research and accelerate progress on robust desktop CUAs. The full dataset has been made public on https://huggingface.co/datasets/vyokky/GUI-360.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。