arXiv:2509.24002cs.CLcs.AI2025-09被引 32

构建首个全面评估大模型外部交互能力的基准测试

MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use

  • 设计127个专家与AI协作的任务,覆盖真实复杂工作流
  • 顶尖模型仅达52.56%单次通过率,平均需16.2步执行
  • 适合评估智能体在真实场景中工具调用与环境交互能力

MCP规范了大语言模型与外部系统交互的方式,是通用智能体的基础。然而现有MCP基准测试范围狭窄,多聚焦于读取密集型任务或交互深度有限的任务,无法体现真实工作流的复杂性与真实性。为此,我们提出MCPMark,一个旨在更真实、全面评估MCP使用的基准测试。该基准包含127个由领域专家与AI协同创建的高质量任务,每个任务均具备定制的初始状态和可程序化验证的脚本。任务要求模型进行丰富多样的环境交互,涵盖广泛的创建、读取、更新、删除(CRUD)操作。我们采用最小代理框架,在工具调用循环中对前沿大模型进行全面评估。实验结果表明,表现最佳的模型gpt-5-medium仅达到52.56% pass@1和33.86% pass^4,而其他广泛认可的强大模型如claude-sonnet-4和o3,其pass@1低于30%,pass^4低于15%。平均而言,大模型每任务需16.2次执行轮次和17.4次工具调用,显著高于以往基准,凸显MCPMark的压测特性。

原文摘要 · Abstract (English)

MCP standardizes how LLMs interact with external systems, forming the foundation for general agents. However, existing MCP benchmarks remain narrow in scope: they focus on read-heavy tasks or tasks with limited interaction depth, and fail to capture the complexity and realism of real-world workflows. To address this gap, we propose MCPMark, a benchmark designed to evaluate MCP use in a more realistic and comprehensive manner. It consists of $127$ high-quality tasks collaboratively created by domain experts and AI agents. Each task begins with a curated initial state and includes a programmatic script for automatic verification. These tasks demand richer and more diverse interactions with the environment, involving a broad range of create, read, update, and delete (CRUD) operations. We conduct a comprehensive evaluation of cutting-edge LLMs using a minimal agent framework that operates in a tool-calling loop. Empirical results show that the best-performing model, gpt-5-medium, reaches only $52.56$\% pass@1 and $33.86$\% pass^4, while other widely regarded strong models, including claude-sonnet-4 and o3, fall below $30$\% pass@1 and $15$\% pass^4. On average, LLMs require $16.2$ execution turns and $17.4$ tool calls per task, significantly surpassing those in previous MCP benchmarks and highlighting the stress-testing nature of MCPMark.

大模型评估智能体工具调用基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。