开源计算机操作智能体框架,推动可复现研究与安全评估
OpenCUA: Open Foundations for Computer-Use Agents
- 构建全流程开源系统,自动采集人类操作数据并生成带推理链的状态-动作对
- 发布首个跨3系统200+应用的大型数据集,72B模型在OSWorld上达45.0%成功率
- 适合关注AI代理安全性、可解释性及开放研究的研究者使用
视觉语言模型在作为计算机操作智能体(CUA)方面展现出强大能力,能自动化执行多样计算机任务。随着其商业潜力增长,最先进系统的具体细节仍为闭源。由于这些智能体将越来越多地代表我们进行数字交互并做出重要决策,研究社区亟需开放的CUA框架来探究其能力、局限与风险。为此,我们提出OpenCUA,一个用于扩展CUA数据与基础模型的综合性开源框架。该框架包含:(1) 无缝捕捉人类计算机操作示范的标注基础设施;(2) AgentNet,首个覆盖3个操作系统和200+应用/网站的大规模计算机操作任务数据集;(3) 可扩展的流水线,将示范转化为带有反思式长链思维推理的状态-动作对,在数据量增加时仍保持性能提升。端到端智能体模型在多个CUA基准测试中表现优异。特别地,OpenCUA-72B在OSWorld-Verified上平均成功率达到45.0%,成为当前开源模型中的最佳水平。进一步分析表明,该方法在不同领域具有良好的泛化能力,且显著受益于增加的推理计算资源。我们公开了标注工具、数据集、代码与模型,以建立开放的CUA研究基础。
原文摘要 · Abstract (English)
Vision-language models have demonstrated impressive capabilities as computer-use agents (CUAs) capable of automating diverse computer tasks. As their commercial potential grows, critical details of the most capable CUA systems remain closed. As these agents will increasingly mediate digital interactions and execute consequential decisions on our behalf, the research community needs access to open CUA frameworks to study their capabilities, limitations, and risks. To bridge this gap, we propose OpenCUA, a comprehensive open-source framework for scaling CUA data and foundation models. Our framework consists of: (1) an annotation infrastructure that seamlessly captures human computer-use demonstrations; (2) AgentNet, the first large-scale computer-use task dataset spanning 3 operating systems and 200+ applications and websites; (3) a scalable pipeline that transforms demonstrations into state-action pairs with reflective long Chain-of-Thought reasoning that sustain robust performance gains as data scales. Our end-to-end agent models demonstrate strong performance across CUA benchmarks. In particular, OpenCUA-72B achieves an average success rate of 45.0% on OSWorld-Verified, establishing a new state-of-the-art (SOTA) among open-source models. Further analysis confirms that our approach generalizes well across domains and benefits significantly from increased test-time computation. We release our annotation tool, datasets, code, and models to build open foundations for further CUA research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。