用专家操作数据训练桌面任务智能体,效果优于以往方法。
Grounding Computer Use Agents on Human Demonstrations
- 基于专家演示构建大规模标注数据集
- 3.56亿标注支持高精度指令-界面映射
- 小样本训练即达顶尖性能,适合通用桌面助手开发
构建可靠的计算机操作智能体需实现语义与屏幕元素的精准对齐。尽管网页和移动应用已有大量数据集,但桌面环境的高质量资源仍稀缺。为此,我们提出GroundCUA,一个基于专家操作演示构建的大规模桌面环境数据集,涵盖12类共87个应用,包含56,000张截图,每张图中所有界面元素均经人工验证标注,总计超过356万条高质量标注。基于这些演示,生成多样化指令以覆盖真实场景任务,为模型训练提供优质数据。利用GroundCUA,我们开发了GroundNext系列模型,通过监督微调在五个基准上达到当前最优表现,且训练数据不足此前工作的十分之一。强化学习后处理进一步提升性能,在OSWorld基准上以o3为规划器的代理设置下,表现可媲美甚至超越使用更多数据训练的模型。结果表明,高质量、专家驱动的数据集对推动通用计算机操作智能体发展至关重要。
原文摘要 · Abstract (English)
Building reliable computer-use agents requires grounding: accurately connecting natural language instructions to the correct on-screen elements. While large datasets exist for web and mobile interactions, high-quality resources for desktop environments are limited. To address this gap, we introduce GroundCUA, a large-scale desktop grounding dataset built from expert human demonstrations. It covers 87 applications across 12 categories and includes 56K screenshots, with every on-screen element carefully annotated for a total of over 3.56M human-verified annotations. From these demonstrations, we generate diverse instructions that capture a wide range of real-world tasks, providing high-quality data for model training. Using GroundCUA, we develop the GroundNext family of models that map instructions to their target UI elements. At both 3B and 7B scales, GroundNext achieves state-of-the-art results across five benchmarks using supervised fine-tuning, while requiring less than one-tenth the training data of prior work. Reinforcement learning post-training further improves performance, and when evaluated in an agentic setting on the OSWorld benchmark using o3 as planner, GroundNext attains comparable or superior results to models trained with substantially more data,. These results demonstrate the critical role of high-quality, expert-driven datasets in advancing general-purpose computer-use agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。