arXiv:2508.20453cs.CL2025-08被引 102

评测大模型在真实复杂任务中使用工具的能力,支持跨工具协同与多步规划。

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

  • 基于MCP协议连接28个真实服务,涵盖250个工具
  • 测试模型从模糊指令中识别工具、规划多步流程的能力
  • 适合评估智能体在金融、科研等跨域任务中的综合表现

我们提出MCP-Bench,一个用于评估大语言模型在需要工具使用、跨工具协调、精确参数控制及规划推理的复杂真实任务中的基准。该基准基于模型上下文协议(MCP),连接了28个代表性的实时MCP服务器,覆盖金融、旅行、科学计算和学术搜索等领域共250个工具。与以往基于API的基准不同,每个MCP服务器提供一组可协同工作的互补工具,支持构建具有丰富输入输出关联的真实多步任务。MCP-Bench测试智能体从模糊指令中检索相关工具、规划多跳执行路径、根据中间工具输出校准响应,并协调跨领域工作流的能力——这些能力在现有依赖显式工具说明、浅层几步流程和孤立领域操作的基准中未被充分评估。我们提出了一个包含工具级模式理解与使用、轨迹级规划和任务完成度的多维度评估框架。对20个先进大模型的实验显示其在该基准上仍存在持续挑战。代码与数据:https://github.com/Accenture/mcp-bench。

原文摘要 · Abstract (English)

We introduce MCP-Bench, a benchmark for evaluating large language models (LLMs) on realistic, multi-step tasks that demand tool use, cross-tool coordination, precise parameter control, and planning/reasoning for solving tasks. Built on the Model Context Protocol (MCP), MCP-Bench connects LLMs to 28 representative live MCP servers spanning 250 tools across domains such as finance, traveling, scientific computing, and academic search. Unlike prior API-based benchmarks, each MCP server provides a set of complementary tools designed to work together, enabling the construction of authentic, multi-step tasks with rich input-output coupling. Tasks in MCP-Bench test agents' ability to retrieve relevant tools from fuzzy instructions without explicit tool names, plan multi-hop execution trajectories for complex objectives, ground responses in intermediate tool outputs, and orchestrate cross-domain workflows - capabilities not adequately evaluated by existing benchmarks that rely on explicit tool specifications, shallow few-step workflows, and isolated domain operations. We propose a multi-faceted evaluation framework covering tool-level schema understanding and usage, trajectory-level planning, and task completion. Experiments on 20 advanced LLMs reveal persistent challenges in MCP-Bench. Code and data: https://github.com/Accenture/mcp-bench.

大模型评估工具使用多步推理跨域协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。