首个评估计算机使用智能体工具调用能力的基准,揭示当前模型工具使用率不足40%。
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
- 构建自动化代码生成流水线,集成158个高质工具覆盖7类常见应用
- 实测显示工具调用使任务成功率从8.3%提升至20.4%(o3模型)
- 首次系统评测工具调用能力,适合研究多模态智能体与人机协同的学者
随着决策与推理能力的进步,多模态智能体在计算机应用场景中展现出巨大潜力。以往评估主要关注图形界面(GUI)交互技能,而由模型上下文协议(MCP)支持的工具调用能力却长期被忽视。仅评估GUI交互的评测方式对集成工具调用的智能体存在天然不公平。本文提出OSWorld-MCP,首个全面且公平的基准,用于评估智能体在真实环境中进行工具调用、GUI操作及决策的能力。我们设计了新颖的自动化代码生成流水线,创建新工具并整合现有工具;经严格人工验证,共获得158个高质量工具(涵盖7类常见应用),每项均通过功能正确性、实用性与通用性检验。对顶尖多模态智能体在该基准上的广泛测试表明,引入MCP工具显著提升任务成功率(如OpenAI o3模型在15步内成功率从8.3%升至20.4%,Claude 4 Sonnet在50步内从40.1%升至43.3%),凸显评估工具调用能力的重要性。然而,即使最强模型工具调用率也仅为36.3%,表明仍有巨大提升空间,彰显该基准的挑战性。通过显式衡量MCP工具使用能力,OSWorld-MCP深化了对多模态智能体的理解,并为复杂、工具辅助环境下的性能评估树立新标准。代码、环境与数据已公开于https://osworld-mcp.github.io。
原文摘要 · Abstract (English)
With advances in decision-making and reasoning capabilities, multimodal agents show strong potential in computer application scenarios. Past evaluations have mainly assessed GUI interaction skills, while tool invocation abilities, such as those enabled by the Model Context Protocol (MCP), have been largely overlooked. Comparing agents with integrated tool invocation to those evaluated only on GUI interaction is inherently unfair. We present OSWorld-MCP, the first comprehensive and fair benchmark for assessing computer-use agents' tool invocation, GUI operation, and decision-making abilities in a real-world environment. We design a novel automated code-generation pipeline to create tools and combine them with a curated selection from existing tools. Rigorous manual validation yields 158 high-quality tools (covering 7 common applications), each verified for correct functionality, practical applicability, and versatility. Extensive evaluations of state-of-the-art multimodal agents on OSWorld-MCP show that MCP tools generally improve task success rates (e.g., from 8.3% to 20.4% for OpenAI o3 at 15 steps, from 40.1% to 43.3% for Claude 4 Sonnet at 50 steps), underscoring the importance of assessing tool invocation capabilities. However, even the strongest models have relatively low tool invocation rates, Only 36.3%, indicating room for improvement and highlighting the benchmark's challenge. By explicitly measuring MCP tool usage skills, OSWorld-MCP deepens understanding of multimodal agents and sets a new standard for evaluating performance in complex, tool-assisted environments. Our code, environment, and data are publicly available at https://osworld-mcp.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。