构建首个面向桌面多模态导航的自动化评测基准,测试智能体从点击到跨应用操作的能力。
OSUniverse: Benchmark for Multimodal GUI-navigation AI Agents
- 按复杂度分级设计任务,涵盖单步点击到多步骤跨应用操作
- 当前最先进智能体表现低于50%,普通白领可完美完成所有任务
- 支持人工与自动评分,自动化验证误差率<2%,适合长期评估
本文提出OSUniverse:一个面向高级图形界面导航人工智能代理的复杂、多模态桌面任务基准,强调易用性、可扩展性、测试案例全面覆盖和自动化验证。任务按复杂度递增划分,从基础精准点击到需技巧、精度和清晰思维的多步骤、多应用测试。第一版基准中,我们校准了测试案例难度,确保发表时最先进(SOTA)代理性能不超过50%,而普通白领可100%准确完成所有任务。基准支持手动评分,同时引入自动化验证机制,平均错误率低于2%。该基准为短期至中期内全自动衡量GUI导航智能体进展、能力与有效性提供了坚实基础。源代码已公开于https://github.com/agentsea/osuniverse。
原文摘要 · Abstract (English)
In this paper, we introduce OSUniverse: a benchmark of complex, multimodal desktop-oriented tasks for advanced GUI-navigation AI agents that focuses on ease of use, extensibility, comprehensive coverage of test cases, and automated validation. We divide the tasks in increasing levels of complexity, from basic precision clicking to multistep, multiapplication tests requiring dexterity, precision, and clear thinking from the agent. In version one of the benchmark, presented here, we have calibrated the complexity of the benchmark test cases to ensure that the SOTA (State of the Art) agents (at the time of publication) do not achieve results higher than 50%, while the average white collar worker can perform all these tasks with perfect accuracy. The benchmark can be scored manually, but we also introduce an automated validation mechanism that has an average error rate less than 2%. Therefore, this benchmark presents solid ground for fully automated measuring of progress, capabilities and the effectiveness of GUI-navigation AI agents over the short and medium-term horizon. The source code of the benchmark is available at https://github.com/agentsea/osuniverse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。