测试101个真实任务,发现大模型用多工具协作成功率不足60%。
LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- 设计动态基准,模拟真实场景下多MCP工具协同调用。
- 实测前沿大模型在复杂任务中成功率达不到60%。
- 揭示7类失败模式,为改进智能体提供具体方向。
工具调用已成为智能体的关键能力。与依赖静态、供应商特定工具定义的传统框架不同,模型上下文协议(MCP)提供统一接口,实现工具的动态发现与调用。然而,现有基准在真实动态场景中对多步骤任务和多样化MCP工具的评估仍存在显著空白。本文提出LiveMCP-101,包含101个需协调使用多个MCP工具的真实世界查询。为应对现实工具响应的时间变异性,引入并行评估框架,让参考智能体同步执行经验证的计划以生成实时参考输出。实验表明,即使前沿大模型在该基准上的成功率也低于60%,暴露出多步工具使用中的关键挑战。全面错误分析识别出七类失败模式,涵盖工具规划、参数设置与输出处理等环节,指明了当前模型改进的具体方向。LiveMCP-101为评估真实世界智能体能力设立了严格标准,推动基于MCP工具编排的自主智能体系统发展。
原文摘要 · Abstract (English)
Tool calling has emerged as a critical capability for AI agents. In contrast to conventional tool calling frameworks that rely on static, provider-specific tool definitions, the Model Context Protocol (MCP) offers a unified interface to discover and invoke tools dynamically. However, there is a significant gap in benchmarking multi-step tasks using diverse MCP tools in realistic, dynamic scenarios. In this work, we present LiveMCP-101, a benchmark of 101 real-world queries that require coordinated use of multiple MCP tools. To address temporal variability in real-world tool responses, we introduce a parallel evaluation framework where a reference agent executes a validated plan simultaneously to produce real-time reference outputs. Experiments show that even frontier LLMs achieve a success rate below 60\%, highlighting challenges in multi-step tool use. Comprehensive error analysis identifies seven failure modes spanning tool planning, parameterization, and output handling, pointing to concrete directions for improving current models. LiveMCP-101 sets a rigorous standard for evaluating real-world agent capabilities, advancing toward autonomous agent systems that reliably execute complex tasks through MCP tool orchestration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。