构建超大规模桌面操作视频数据集,助力通用计算机代理发展
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
- 收集87种应用的连续30帧视频与光标轨迹,含600万帧专家操作数据
- 提供55小时高质量视频和3.6万张带标注界面元素的截图,支持多任务训练
- 适合作为复杂桌面操作研究的基准,尤其适合视觉-动作联合建模方向
计算机使用代理(CUAs)有望自动化复杂的桌面工作流,但通用化进展受限于高质量连续视频演示数据的匮乏。现有最大公开数据集ScaleCUA仅含200万张静态截图,不足20小时视频。为此,我们推出CUA-Suite,一个面向专业桌面代理的大型视频演示生态。核心是VideoCUA,涵盖87种应用的约1万项人类示范任务,提供连续30帧/秒屏幕录制、光标运动轨迹及多层次推理标注,总计约55小时、600万帧专家视频。相比仅记录最终点击坐标的稀疏数据,连续视频完整保留了人机交互的时间动态,可无损转换为现有代理框架所需格式。CUA-Suite还包含两个互补资源:UI-Vision,用于评估代理的视觉定位与规划能力;GroundCUA,一个包含5.6万张标注截图和超过360万条界面元素标注的大规模定位数据集。初步评估显示,当前基础动作模型在专业桌面应用中失败率高达60%以上。此外,该数据集支持通用屏幕解析、连续空间控制、基于视频的奖励建模及视觉世界模型等新兴方向。所有数据与模型均已开源。
原文摘要 · Abstract (English)
Computer-use agents (CUAs) hold great promise for automating complex desktop workflows, yet progress toward general-purpose agents is bottlenecked by the scarcity of continuous, high-quality human demonstration videos. Recent work emphasizes that continuous video, not sparse screenshots, is the critical missing ingredient for scaling these agents. However, the largest existing open dataset, ScaleCUA, contains only 2 million screenshots, equating to less than 20 hours of video. To address this bottleneck, we introduce CUA-Suite, a large-scale ecosystem of expert video demonstrations and dense annotations for professional desktop computer-use agents. At its core is VideoCUA, which provides approximately 10,000 human-demonstrated tasks across 87 diverse applications with continuous 30 fps screen recordings, kinematic cursor traces, and multi-layerfed reasoning annotations, totaling approximately 55 hours and 6 million frames of expert video. Unlike sparse datasets that capture only final click coordinates, these continuous video streams preserve the full temporal dynamics of human interaction, forming a superset of information that can be losslessly transformed into the formats required by existing agent frameworks. CUA-Suite further provides two complementary resources: UI-Vision, a rigorous benchmark for evaluating grounding and planning capabilities in CUAs, and GroundCUA, a large-scale grounding dataset with 56K annotated screenshots and over 3.6 million UI element annotations. Preliminary evaluation reveals that current foundation action models struggle substantially with professional desktop applications (~60% task failure rate). Beyond evaluation, CUA-Suite's rich multimodal corpus supports emerging research directions including generalist screen parsing, continuous spatial control, video-based reward modeling, and visual world models. All data and models are publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。