用网络视频自动学习人类操作电脑,提升智能助手的实用能力。
Watch and Learn: Learning to Use Computers from Online Videos
- 将视频画面变化逆向推导为操作动作,无需人工标注。
- 生成超5.3万条高质量操作轨迹,显著提升智能体表现。
- 适合想训练真实场景下电脑操作智能体的研究者使用。
计算机使用智能体(CUAs)需在多样且不断演进的应用中规划任务流程,但受限于缺乏大规模高质量训练数据。现有数据集范围窄、静态且标注成本高,而合成数据常导致行为过于简化或不准确。我们提出 Watch & Learn(W&L)框架,可将互联网上现成的人类电脑操作视频规模化转化为可执行的界面轨迹。不同于直接生成动作或依赖手工规则,我们将轨迹标注建模为逆动力学问题:从连续屏幕状态预测用户操作,简化学习并提升跨领域泛化能力。通过任务感知的检索与标注流程,W&L生成超过53,000条高质量轨迹,既可作为上下文示例,也可用于监督训练。在OSWorld上,它持续提升通用与专用CUAs性能;在WindowsAgentArena上,70亿参数模型在15步限制内达到当前最佳表现。结果表明,网络规模的人类示范视频可成为推进真实世界CUAs的实用且可扩展基础。
原文摘要 · Abstract (English)
Computer-using agents (CUAs) must plan task workflows across diverse and evolving applications, yet progress is limited by the lack of large-scale, high-quality training data. Existing datasets are narrow, static, and costly to annotate, while synthetic data often yields oversimplified or misaligned behaviors. We present Watch & Learn (W&L), a framework that converts readily available Internet videos of human computer use into executable UI trajectories at scale. Instead of directly generating actions or relying on handcrafted heuristics, we cast trajectory annotation as an inverse dynamics problem that predicts user actions from consecutive screen states, which simplifies learning and generalizes across domains. Through a task-aware retrieval and labeling pipeline, W&L yields over 53K high-quality trajectories that enhance CUAs both as in-context exemplars and as supervised training data. On OSWorld, it consistently improves general-purpose and specialized CUAs, while on WindowsAgentArena it achieves state-of-the-art performance among 7B-scale models under the 15-step limit. These results show that web-scale human demonstration videos can serve as a practical and scalable foundation for advancing real-world CUAs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。