用自动化框架评估大模型智能体在多个领域的表现。
MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models
- 基于MCP协议自动构建任务并评估智能体
- 在五个真实场景中验证了评估效果
- 开源工具,适合研究者和开发者使用
基于大语言模型的智能体迅速发展,亟需稳健且可扩展的评估框架。现有方法依赖静态基准和人工数据收集,难以满足实际评估需求。我们提出MCPEval,一个基于模型上下文协议(MCP)的开源框架,可自动化完成跨多领域的任务生成与深度评估。该框架统一评估指标,无缝集成原生智能体工具,无需手动搭建评估流程。在五个真实应用场景中的实证结果表明,其能有效揭示领域特异性性能差异。我们已公开发布MCPEval(https://github.com/SalesforceAIResearch/MCPEval),以推动可复现、标准化的大型语言模型智能体评估。
原文摘要 · Abstract (English)
The rapid rise of Large Language Models (LLMs)-based intelligent agents underscores the need for robust, scalable evaluation frameworks. Existing methods rely on static benchmarks and labor-intensive data collection, limiting practical assessment. We introduce MCPEval, an open-source Model Context Protocol (MCP)-based framework that automates end-to-end task generation and deep evaluation of LLM agents across diverse domains. MCPEval standardizes metrics, seamlessly integrates with native agent tools, and eliminates manual effort in building evaluation pipelines. Empirical results across five real-world domains show its effectiveness in revealing nuanced, domain-specific performance. We publicly release MCPEval https://github.com/SalesforceAIResearch/MCPEval to promote reproducible and standardized LLM agent evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。