构建跨平台开源电脑操作数据集,提升通用电脑代理性能。
ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data
- 用自动化代理与人类专家闭环协作构建跨平台数据集。
- 在多个基准上实现显著提升,最高达94.4%准确率。
- 适合研究通用计算机操作智能体与数据驱动方法的学者。
视觉语言模型(VLMs)已使计算机使用代理(CUAs)能够自主操作图形界面,展现出巨大潜力,但进展受限于缺乏大规模开源电脑使用数据和基础模型。本文提出ScaleCUA,迈向开源电脑使用代理的规模化。该数据集覆盖6个操作系统和3个任务领域,通过自动化代理与人类专家的闭环流程构建。基于该数据集训练的ScaleCUA可在多平台无缝运行,显著优于基线模型:WebArena-Lite-v2上提升26.6分,ScreenSpot-Pro上提升10.7分;并在多个基准上刷新纪录:MMBench-GUI L1-Hard达到94.4%,OSWorld-G达60.6%,WebArena-Lite-v2达47.4%。这些结果验证了数据驱动扩展对通用电脑使用代理的有效性。我们将公开数据、模型与代码,以推动后续研究:https://github.com/OpenGVLab/ScaleCUA。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have enabled computer use agents (CUAs) that operate GUIs autonomously, showing great potential, yet progress is limited by the lack of large-scale, open-source computer use data and foundation models. In this work, we introduce ScaleCUA, a step toward scaling open-source CUAs. It offers a large-scale dataset spanning 6 operating systems and 3 task domains, built via a closed-loop pipeline uniting automated agents with human experts. Trained on this scaled-up data, ScaleCUA can operate seamlessly across platforms. Specifically, it delivers strong gains over baselines (+26.6 on WebArena-Lite-v2, +10.7 on ScreenSpot-Pro) and sets new state-of-the-art results (94.4% on MMBench-GUI L1-Hard, 60.6% on OSWorld-G, 47.4% on WebArena-Lite-v2). These findings underscore the power of data-driven scaling for general-purpose computer use agents. We will release data, models, and code to advance future research: https://github.com/OpenGVLab/ScaleCUA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。