arXiv:2602.23367cs.AIcs.IR2026-02

首个模拟真实用户查询的MCP工具评估数据集,提升工具检索评测真实性。

HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance

  • 基于2800个工具、308个MCP服务器构建多样化用户查询
  • 每工具配多个人设,覆盖精准与模糊请求场景
  • 适合评估LLM工具调用能力,尤其关注真实交互泛化性

模型上下文协议(MCP)服务器整合数千个开源标准化工具,连接大语言模型与外部系统;然而现有数据集和基准缺乏真实、类人用户查询,制约了对MCP工具使用及生态的评估。现有数据常包含工具描述,却无法体现用户表达请求的多样性,导致评估结果泛化性差且可靠性虚高。本文提出首个大规模MCP数据集,基于MCP Zero构建,涵盖308个MCP服务器上的2800个工具,为每个工具生成多个独特用户角色,覆盖从精确任务指令到模糊探索性命令的多种意图层次,真实反映现实交互复杂性。

原文摘要 · Abstract (English)

Model Context Protocol (MCP) servers contain a collection of thousands of open-source standardized tools, linking LLMs to external systems; however, existing datasets and benchmarks lack realistic, human-like user queries, remaining a critical gap in evaluating the tool usage and ecosystems of MCP servers. Existing datasets often do contain tool descriptions but fail to represent how different users portray their requests, leading to poor generalization and inflated reliability of certain benchmarks. This paper introduces the first large-scale MCP dataset featuring diverse, high-quality diverse user queries generated specifically to match 2800 tools across 308 MCP servers, developing on the MCP Zero dataset. Each tool is paired with multiple unique user personas that we have generated, to capture varying levels of user intent ranging from precise task requests, and ambiguous, exploratory commands, reflecting the complexity of real-world interaction patterns.

MCP数据集工具检索评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。