构建五层评测体系,精准评估大模型在工具协同中的真实能力
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
- 通过等效功能集构建真实标注,避免评价偏差
- 实测顶尖代理在跨服务器任务中仍存在系统性弱点
- 适合研究工具调用、多跳推理与智能体鲁棒性的学者
我们提出ETOM,一个五层级基准,用于评估大模型智能体在分层模型-上下文协议(MCP)生态中执行多跳端到端工具协同的能力。现有评测常孤立评估工具,忽略功能重叠与跨服务器协调等挑战,导致评价过于乐观。ETOM通过“等效功能集”构建真实标注,支持客观指标如F1分数,减少对大模型作为评判者的依赖。其五级递进式课程系统测试智能体能力,从单工具协同到复杂跨服务器规划,以及对超出范围请求的鲁棒性。实验表明,僵化层级结构会抑制性能,即使最先进代理也存在系统性脆弱性。ETOM提供诊断框架,揭示这些局限并指导更高效工具使用智能体的发展。
原文摘要 · Abstract (English)
We introduce ETOM, a five-level benchmark for evaluating multi-hop, end-to-end tool orchestration by LLM agents within a hierarchical Model-Context Protocol (MCP) ecosystem. Existing benchmarks often assess tools in isolation, overlooking challenges such as functional overlap and cross-server orchestration, which can lead to overly optimistic evaluations. ETOM addresses these gaps by constructing ground truth through "equal function sets", enabling objective metrics such as F1 score and reducing reliance on LLM-as-a-judge evaluation. Its five-level curriculum systematically tests agent capabilities, from single-tool orchestration to complex cross-server planning, as well as robustness to out-of-scope requests. Experiments reveal that rigid hierarchies can hinder performance without co-designed strategies, and even state-of-the-art agents exhibit systemic weaknesses in robustness. ETOM provides a diagnostic framework to expose these limitations and guide the development of more capable and efficient tool-using agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。