arXiv:2512.24565cs.AI2025-12被引 10

构建真实MCP工具任务基准,评估大模型工具使用能力。

MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use

  • 基于真实MCP定义构建任务与模拟工具集
  • 动态沙盒中测试模型选工具与辨识能力,区分率差异显著
  • 适合评估多步工具调用的智能体,开源可用

大型语言模型正越来越多地作为自主代理使用,其通过模型上下文协议(MCP)调用外部工具被视为未来趋势。现有MCP评估数据集存在依赖外部MCP服务、缺乏难度感知等问题。为此,我们提出MCPAgentBench,一个基于真实MCP定义的基准,用于评估代理的工具使用能力。我们构建了一个包含真实任务和模拟MCP工具的数据集。评估采用动态沙盒环境,向代理提供含干扰项的候选工具列表,以测试其工具选择与辨识能力。此外,我们引入全面指标,衡量任务完成率与执行效率。在多个主流大模型上的实验表明,不同模型在处理复杂多步工具调用时表现差异显著。所有代码已开源至GitHub。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a future trend. Current MCP evaluation sets suffer from issues such as reliance on external MCP services and a lack of difficulty awareness. To address these limitations, we propose MCPAgentBench, a benchmark based on real-world MCP definitions designed to evaluate the tool-use capabilities of agents. We construct a dataset containing authentic tasks and simulated MCP tools. The evaluation employs a dynamic sandbox environment that presents agents with candidate tool lists containing distractors, thereby testing their tool selection and discrimination abilities. Furthermore, we introduce comprehensive metrics to measure both task completion rates and execution efficiency. Experiments conducted on various latest mainstream Large Language Models reveal significant performance differences in handling complex, multi-step tool invocations. All code is open-source at Github.

大模型代理工具调用评估基准MCP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。