首次系统评估计算机代理效率,发现其耗时是人类的2.7至4.3倍。
OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
- 构建人工标注的OSWorld Human数据集,提供人类执行路径作为基准
- 发现大模型调用导致延迟飙升,后期步骤耗时可达初期3倍
- 当前最佳代理步数超人类必要步数2.7至4.3倍,效率严重不足
生成式AI正被用于解决涉及桌面应用的各类计算机操作任务。当前顶尖系统仅关注在主流基准上的准确率提升,但这些系统因端到端延迟极高(如任务耗时数十分钟,而人类仅需几分钟)而实际不可用。为理解此现象并指导未来计算机代理发展,我们首次在计算机操作人工智能旗舰基准OSWorld上开展时序性能研究。结果发现,大模型用于规划、反思与判断的调用占整体延迟的绝大部分;随着任务步骤增加,每一步耗时可比初始阶段长3倍。随后,我们构建了OSWorld Human——原始OSWorld数据集的人工标注版本,包含每个任务的人类执行轨迹。在该数据集上评估16个代理的效率,发现即使表现最优的代理也需比必要步数多出2.7至4.3倍。
原文摘要 · Abstract (English)
Generative AI is being leveraged to solve a variety of computer-use tasks involving desktop applications. State-of-the-art systems have focused solely on improving accuracy on leading benchmarks. However, these systems are practically unusable due to extremely high end-to-end latency (e.g., tens of minutes) for tasks that typically take humans just a few minutes to complete. To understand the cause behind this and to guide future developments of computer agents, we conduct the first study on the temporal performance of computer-use agents on OSWorld, the flagship benchmark in computer-use AI. We find that large model calls for planning, reflection, and judging account for most of the overall latency, and as an agent uses more steps to complete a task, each successive step can take 3x longer than steps at the beginning of a task. We then construct OSWorld Human, a manually annotated version of the original OSWorld dataset that contains a human-determined trajectory for each task. We evaluate 16 agents on their efficiency using OSWorld Human and found that even the best agents take 2.7-4.3x more steps than necessary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。