arXiv:2504.11094cs.IRcs.DB2025-04被引 17

评测主流MCP服务器性能,发现声明式接口能显著提升准确率。

Evaluation Report on MCP Servers

  • 构建MCPBench评估框架,对比多个MCP服务器的性能表现。
  • 最佳方案Bing Web Search准确率达64%,但整体效率仍有提升空间。
  • 声明式接口可大幅提升准确率,适合关注AI数据检索优化的研究者。

随着大模型的发展,自2024年底以来涌现出大量模型上下文协议(MCP)服务。然而,MCP服务器的有效性与效率尚未得到充分研究。为此,我们提出了一个名为MCPBench的评估框架,选取多个广泛使用的MCP服务器,在准确性、耗时和令牌使用量方面进行了实验评估。实验结果表明,表现最优异的MCP服务——Bing Web Search,准确率达到64%。值得注意的是,我们发现引入声明式接口可显著提升MCP服务器的准确性。该研究为后续优化MCP实现提供了基础,有助于推动更高效的AI驱动应用与数据检索解决方案的发展。

原文摘要 · Abstract (English)

With the rise of LLMs, a large number of Model Context Protocol (MCP) services have emerged since the end of 2024. However, the effectiveness and efficiency of MCP servers have not been well studied. To study these questions, we propose an evaluation framework, called MCPBench. We selected several widely used MCP server and conducted an experimental evaluation on their accuracy, time, and token usage. Our experiments showed that the most effective MCP, Bing Web Search, achieved an accuracy of 64%. Importantly, we found that the accuracy of MCP servers can be substantially enhanced by involving declarative interface. This research paves the way for further investigations into optimized MCP implementations, ultimately leading to better AI-driven applications and data retrieval solutions.

MCP评估检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。