arXiv:2409.08264cs.AI2024-09ICML被引 199

构建可扩展的Windows操作系统代理评测环境,支持多模态任务自动化。

Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

论文配图:Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
图 1 · 摘自论文原文
  • 基于真实Windows系统构建可复现的多模态代理评测环境。
  • 150+任务覆盖多个领域,全量评估仅需20分钟。
  • 新提出的Navi代理在复杂任务中达19.5%成功率,适合研究智能体与人机协作。

大型语言模型(LLMs)在作为计算机代理方面展现出巨大潜力,能提升人类在需要规划与推理的多模态任务中的生产力和软件可访问性。然而,由于现有基准测试通常局限于特定模态或领域(如纯文本、网页导航、问答、编程),且完整评估因任务的多步序列特性耗时数天,导致真实环境中代理性能的衡量仍具挑战。为此,我们提出Windows Agent Arena:一个专注于真实Windows操作系统的可复现通用环境,代理可在其中自由使用人类用户常见的各类应用、工具和浏览器完成任务。我们基于OSWorld框架(Xie et al., 2024)构建了150多个跨代表性领域的多样化任务,涵盖规划、屏幕理解与工具使用能力。该基准支持无缝并行化,在Azure上可将全量评估缩短至20分钟以内。为展示其能力,我们引入新型多模态代理Navi,其在Windows域任务中取得19.5%的成功率,相较未辅助的人类74.5%表现仍有差距。同时,Navi在主流网页基准Mind2Web上也表现出色。我们提供了对Navi性能的定量与定性分析,并揭示了未来智能体开发与数据生成的研究机遇。

原文摘要 · Abstract (English)

Large language models (LLMs) show remarkable potential to act as computer agents, enhancing human productivity and software accessibility in multi-modal tasks that require planning and reasoning. However, measuring agent performance in realistic environments remains a challenge since: (i) most benchmarks are limited to specific modalities or domains (e.g. text-only, web navigation, Q&A, coding) and (ii) full benchmark evaluations are slow (on order of magnitude of days) given the multi-step sequential nature of tasks. To address these challenges, we introduce the Windows Agent Arena: a reproducible, general environment focusing exclusively on the Windows operating system (OS) where agents can operate freely within a real Windows OS and use the same wide range of applications, tools, and web browsers available to human users when solving tasks. We adapt the OSWorld framework (Xie et al., 2024) to create 150+ diverse Windows tasks across representative domains that require agent abilities in planning, screen understanding, and tool usage. Our benchmark is scalable and can be seamlessly parallelized in Azure for a full benchmark evaluation in as little as 20 minutes. To demonstrate Windows Agent Arena's capabilities, we also introduce a new multi-modal agent, Navi. Our agent achieves a success rate of 19.5% in the Windows domain, compared to 74.5% performance of an unassisted human. Navi also demonstrates strong performance on another popular web-based benchmark, Mind2Web. We offer extensive quantitative and qualitative analysis of Navi's performance, and provide insights into the opportunities for future research in agent development and data generation using Windows Agent Arena. Webpage: https://microsoft.github.io/WindowsAgentArena Code: https://github.com/microsoft/WindowsAgentArena

智能体评测多模态操作系统自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。