arXiv:2509.09734cs.CLcs.AI2025-09AAAI被引 34

构建MCP工具交互基准,真实评估语言代理能力。

MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools

  • 基于MCP协议搭建33个服务器、188个工具的测试环境。
  • 设计600个跨6类复杂度的任务,覆盖真实使用场景。
  • 采用结果导向评估法,突出实际任务完成率。

Model Context Protocol(MCP)正迅速成为关键的开放标准,旨在提升智能体与工具的集成与互操作性,有望开启强大、互联且真正实用的智能体式AI新时代。然而,尽管MCP日益普及,现有基准往往无法捕捉该范式下的真实智能体表现,导致对其实际价值的认知失真,难以可靠区分性能差异。为填补这一评估空白,我们提出MCP-AgentBench——一个专为评估MCP驱动工具交互中的语言智能体能力而设计的综合性基准。核心贡献包括:构建由33个运行服务器和188个独立工具组成的稳健MCP测试平台;开发包含600个系统设计查询、覆盖6种不同交互复杂度类别的基准任务集;引入MCP-Eval——一种以任务成果为导向的新型评估方法,强调真实世界任务的成功。通过对领先语言智能体的广泛实证评估,我们提供了基础性洞见。MCP-AgentBench旨在为研究社区提供标准化、可靠的框架,助力构建、验证并推进能够充分释放MCP变革潜力的智能体,从而加速实现真正强大且可互操作的AI系统。

原文摘要 · Abstract (English)

The Model Context Protocol (MCP) is rapidly emerging as a pivotal open standard, designed to enhance agent-tool integration and interoperability, and is positioned to unlock a new era of powerful, interconnected, and genuinely utilitarian agentic AI. However, despite MCP's growing adoption, existing benchmarks often fail to capture real-world agent performance within this new paradigm, leading to a distorted perception of their true operational value and an inability to reliably differentiate proficiencies. To bridge this critical evaluation gap, we introduce MCP-AgentBench -- a comprehensive benchmark specifically engineered to rigorously assess language agent capabilities in MCP-mediated tool interactions. Core contributions of MCP-AgentBench include: the establishment of a robust MCP testbed comprising 33 operational servers with 188 distinct tools; the development of a benchmark featuring 600 systematically designed queries distributed across 6 distinct categories of varying interaction complexity; and the introduction of MCP-Eval, a novel outcome-oriented evaluation methodology prioritizing real-world task success. Through extensive empirical evaluation of leading language agents, we provide foundational insights. MCP-AgentBench aims to equip the research community with a standardized and reliable framework to build, validate, and advance agents capable of fully leveraging MCP's transformative benefits, thereby accelerating progress toward truly capable and interoperable AI systems.

智能体评估MCP协议工具交互基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。