arXiv:2505.16700cs.AI2025-05被引 50

首个评估大模型工具调用能力的多维基准,覆盖6大场景507个任务。

MCP-RADAR: A Multi-Dimensional Benchmark for Evaluating Tool Use Capabilities in Large Language Models

  • 构建6大领域507个任务的MCP框架评测集,模拟真实工具使用流程。
  • 量化评估答案正确率与操作准确率,包含调用次数和资源效率等指标。
  • 开源完整工具链与数据,助力开发者优化大模型代理性能。

随着大语言模型(LLMs)从被动文本生成转向可与外部工具交互的主动推理代理,模型上下文协议(MCP)已成为动态工具发现与编排的关键标准化框架。尽管其在工业界广泛应用,现有评估方法未能充分衡量该范式下的工具使用能力。本文提出MCP-RADAR,首个专为MCP框架设计的综合性评测基准。该基准包含涵盖数学推理、网络搜索、邮件、日历、文件管理和终端操作六大领域的507个挑战性任务,以答案正确性和操作准确性为主要评估标准。评估采用真实MCP工具及高保真官方工具仿真,实现对计算资源效率和成功工具调用轮次等客观量化指标的测量。相较于依赖人工判断或二值成功标准的传统基准,MCP-RADAR提供跨多任务域的客观评估。对主流闭源与开源大模型的测评揭示了显著的能力差异,并指出准确率与效率间的权衡关系。研究结果为大模型开发者与工具创建者提供可操作洞见,确立了适用于更广泛大模型智能体生态的标准化方法。所有实现、配置与数据集已公开于https://anonymous.4open.science/r/MCPRadar-B143。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) evolve from passive text generators to active reasoning agents capable of interacting with external tools, the Model Context Protocol (MCP) has emerged as a key standardized framework for dynamic tool discovery and orchestration. Despite its widespread industry adoption, existing evaluation methods do not adequately assess tool utilization capabilities under this new paradigm. To address this gap, this paper introduces MCP-RADAR, the first comprehensive benchmark specifically designed to evaluate LLM performance within the MCP framework. MCP-RADAR features a challenging dataset of 507 tasks spanning six domains: mathematical reasoning, web search, email, calendar, file management, and terminal operations. It quantifies performance based on two primary criteria: answer correctness and operational accuracy. To closely emulate real-world usage, our evaluation employs both authentic MCP tools and high-fidelity simulations of official tools. Unlike traditional benchmarks that rely on subjective human evaluation or binary success metrics, MCP-RADAR adopts objective, quantifiable measurements across multiple task domains, including computational resource efficiency and the number of successful tool-invocation rounds. Our evaluation of leading closed-source and open-source LLMs reveals distinct capability profiles and highlights a significant trade-off between accuracy and efficiency. Our findings provide actionable insights for both LLM developers and tool creators, establishing a standardized methodology applicable to the broader LLM agent ecosystem. All implementations, configurations, and datasets are publicly available at https://anonymous.4open.science/r/MCPRadar-B143.

大模型评估工具调用MCP基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。