arXiv:2508.12566cs.AI2025-08被引 5

首次系统评估大模型用外部工具的能力,发现实际效果远低于预期。

Help or Hurdle? Rethinking Model Context Protocol-Augmented Large Language Models

  • 构建四维评估框架,测试模型主动调用工具、遵守指令等行为。
  • 在20,000次API调用中发现,多数模型工具使用效率低且成本高。
  • 为开发可控的工具增强模型提供可复现的基准,适合研究者参考。

模型上下文协议(MCP)使大语言模型(LLMs)能按需访问外部资源。尽管普遍认为其可提升性能,但模型如何真正利用这一能力仍不明确。本文提出MCPGAUGE,首个全面评估LLM-MCP交互的框架,涵盖四大维度:主动性(自主调用工具)、合规性(遵循指令)、有效性(任务完成度)和开销(计算成本)。该框架包含160个提示和25个数据集,覆盖知识理解、通用推理与代码生成。大规模实验涉及六款商业大模型、30套MCP工具,共发起约20,000次API调用,计算成本超6,000美元。研究揭示四项挑战主流认知的发现,指出现有AI-工具融合的关键局限,并将MCPGAUGE定位为推动可控、工具增强型大模型发展的原则性基准。

原文摘要 · Abstract (English)

The Model Context Protocol (MCP) enables large language models (LLMs) to access external resources on demand. While commonly assumed to enhance performance, how LLMs actually leverage this capability remains poorly understood. We introduce MCPGAUGE, the first comprehensive evaluation framework for probing LLM-MCP interactions along four key dimensions: proactivity (self-initiated tool use), compliance (adherence to tool-use instructions), effectiveness (task performance post-integration), and overhead (computational cost incurred). MCPGAUGE comprises a 160-prompt suite and 25 datasets spanning knowledge comprehension, general reasoning, and code generation. Our large-scale evaluation, spanning six commercial LLMs, 30 MCP tool suites, and both one- and two-turn interaction settings, comprises around 20,000 API calls and over USD 6,000 in computational cost. This comprehensive study reveals four key findings that challenge prevailing assumptions about the effectiveness of MCP integration. These insights highlight critical limitations in current AI-tool integration and position MCPGAUGE as a principled benchmark for advancing controllable, tool-augmented LLMs.

大模型工具调用评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。