测试大模型在不断变化的工具接口中完成任务的能力,发现顶级模型也难适应。
MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers

- 设计11种变异操作模拟工具接口真实演化,覆盖123个MCP服务器
- 12个主流大模型在演化后接口上性能下降,最高达14.4%
- 适合关注大模型工具使用鲁棒性的研究者和开发者
随着模型上下文协议(MCP)服务器成为连接大语言模型与外部工具的核心基础设施,现有评估基准虽使用真实MCP服务器测试大模型代理的工具调用能力,却忽略了工具接口和功能随时间持续演化的现实。这导致评估结果无法反映代理在动态工具环境中的适应能力。为此,我们提出MCPEvol-Bench,一个用于评估大模型代理在动态工具集演化下任务求解能力的新基准。基于大规模实证研究,我们设计了11种变异操作,模拟123个MCP服务器中的真实工具演化。我们在多个版本的MCP服务器上对12个前沿大模型进行评测,发现即使最先进的模型也难以适应演化后的工具:例如,GPT-5.4和Claude-Sonnet-4-6在演化后的服务器上性能分别下降13.7%和14.4%,同时规划与推理错误显著增加。这些发现揭示了大模型工作流在动态环境中的脆弱性,确立了MCPEvol-Bench作为评估代理适应能力的标准。
原文摘要 · Abstract (English)
As Model Context Protocol (MCP) servers emerge as the core infrastructure for connecting LLMs with external tools, existing benchmarks leverage real-world MCP servers to evaluate LLM agents' tool-using capabilities. However, these benchmarks overlook the continuous evolution of tool interfaces and functionalities within MCP servers, resulting in flawed assessments that fail to capture the agent's adaptability in changing tool landscapes. To bridge this gap, we introduce \textbf{MCPEvol-Bench}, a novel benchmark for evaluating the task-solving capabilities of LLM agents under dynamic toolset evolution. Inspired by large-scale empirical study, we propose 11 mutation operators to simulate realistic tool evolution within 123 MCP servers. We benchmark 12 state-of-the-art LLMs on multiple versions of MCP servers, revealing that even frontier models struggle to adapt to evolving tools. For instance, GPT-5.4 and Claude-Sonnet-4-6 exhibit performance declines of 13.7\% and 14.4\% in evolved MCP servers, respectively, accompanied by substantial increases in planning and reasoning errors. These findings highlight the vulnerability of LLM-driven workflows, establishing MCPEvol-Bench as a standard for evaluating agent adaptability in dynamic tool environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。