arXiv:2508.16260cs.CLcs.AI2025-08被引 13

构建超大规模真实工具评测集,推动大模型智能体能力评估

MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use

  • 引入550+可执行真实工具,构建超14万token动作空间
  • 实测显示多数模型在复杂工具环境中性能下降,但智能体模型表现更优
  • 适合研究智能体、工具调用与真实场景评估的开发者与研究人员

大型语言模型正从文本生成器演变为具备推理能力的智能体,其使用外部工具的能力成为关键。然而,现有评测基准受限于合成工具和狭窄的动作空间。为此,我们提出MCPVerse,一个面向真实世界场景的广阔评测基准。该基准整合超过550个真实可执行工具,构建了超过14万令牌的动作空间,并采用基于结果的实时真值评估方法,适用于时效性任务。我们在三种模式(Oracle、Standard、Max-Scale)下对当前最先进的大模型进行评测,发现尽管多数模型在面对更大工具集时性能下降,但如Claude-4-Sonnet等智能体模型能有效利用扩展的探索空间提升准确率。这一发现揭示了主流模型在复杂现实场景中的局限性,同时确立MCPVerse作为衡量与推进智能体工具使用能力的关键基准。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are evolving from text generators into reasoning agents. This transition makes their ability to use external tools a critical capability. However, evaluating this skill presents a significant challenge. Existing benchmarks are often limited by their reliance on synthetic tools and severely constrained action spaces. To address these limitations, we introduce MCPVerse, an expansive, real-world benchmark for evaluating agentic tool use. MCPVerse integrates more than 550 real-world, executable tools to create an unprecedented action space exceeding 140k tokens, and employs outcome-based evaluation with real-time ground truth for time-sensitive tasks. We benchmarked the state-of-the-art LLMs across three modes (Oracle, Standard, and Max-Scale), revealing that while most models suffer performance degradation when confronted with larger tool sets, the agentic models, such as Claude-4-Sonnet, can effectively leverage expanded exploration spaces to improve accuracy. This finding not only exposes the limitations of state-of-the-art models in complex, real-world scenarios but also establishes MCPVerse as a critical benchmark for measuring and advancing agentic tool use capabilities.

智能体工具使用评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。