研究计算机操作代理的不可靠性根源,提出需在重复执行中评估性能。
On the Reliability of Computer Use Agents

- 通过三因素分析:执行随机性、任务描述模糊、代理行为变异
- 发现任务指定方式与行为稳定性共同决定可靠性,重复执行成功率差异显著
- 适合关注自动化系统鲁棒性的研究人员与工程师
计算机使用代理在网页导航、桌面自动化和软件交互等真实任务上快速进步,某些情况下已超越人类表现。然而,即使任务与模型不变,同一代理在一次成功后可能在重复执行中失败。这引发核心问题:为何代理能成功一次却无法可靠重复?本文通过三个因素研究其不可靠性根源:执行过程中的随机性、任务说明的模糊性以及代理行为的变异性。我们在OSWorld数据集上进行多次重复任务执行,并采用配对统计检验捕捉不同设置下的任务级变化。分析表明,可靠性既取决于任务如何指定,也受代理行为跨运行变化的影响。研究建议应在重复执行下评估代理,允许代理通过交互解决任务模糊性,并优先选择跨运行稳定的策略。
原文摘要 · Abstract (English)
Computer-use agents have rapidly improved on real-world tasks such as web navigation, desktop automation, and software interaction, in some cases surpassing human performance. Yet even when the task and model are unchanged, an agent that succeeds once may fail on a repeated execution of the same task. This raises a fundamental question: if an agent can succeed at a task once, what prevents it from doing so reliably? In this work, we study the sources of unreliability in computer-use agents through three factors: stochasticity during execution, ambiguity in task specification, and variability in agent behavior. We analyze these factors on OSWorld using repeated executions of the same task together with paired statistical tests that capture task-level changes across settings. Our analysis shows that reliability depends on both how tasks are specified and how agent behavior varies across executions. These findings suggest the need to evaluate agents under repeated execution, to allow agents to resolve task ambiguity through interaction, and to favor strategies that remain stable across runs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。