用工具说明自动生成真实测评场景,无需人工干预。
Agent Seer: Synthesizing Scenarios from Specification Understanding

- 从工具说明中提取语义信息,自动生成多轮对话
- 在7个不同领域中实现完整工具覆盖和高正确率
- 揭示参数复杂度是质量差异主因,值准确性易出错
评估使用外部工具的AI代理需要能反映实践者工具组合与对话迭代的真实测试场景。手工构建此类场景需深厚领域知识,难以扩展且无法跟踪不断变化的API。我们发现工具规格(函数名、自然语言描述、类型化参数模式)已蕴含足够语义信息,可无须人工标注或实时工具调用,直接合成真实测评场景。Agent Seer基于单个Model Context Protocol (MCP) 规格,无需示例、无需实时工具访问、无需领域调优,通过丰富原始模式、生成带合成输出的分级场景,并扩展为具备真实数据模拟的多轮对话,表现出强工具调用正确性与对话连贯性。在涵盖多个领域和工具集规模的7个MCP规格上评估,结果表明该流程在所有领域均表现良好,小中型规格实现完全工具覆盖。分析发现:参数模式复杂度是质量波动的最强相关因素——工具集大小影响较小且独立;参数值准确性是不完美场景中的主要失败模式,这一子维度对粗粒度名称匹配指标不可见。
原文摘要 · Abstract (English)
Evaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain expertise, does not scale across tool ecosystems, and produces static benchmarks that cannot track evolving APIs. We observe that tool specifications -- function names, natural-language descriptions, and typed parameter schemas -- already encode sufficient semantic information to synthesize realistic evaluation scenarios without manual curation or live tool execution. Agent Seer builds off this latent information: from a single Model Context Protocol (MCP) specification, with no examples, no live tool access, and no domain-specific tuning. This pipeline enriches raw schemas, generates graded scenarios with synthetic tool outputs, and expands them into mock-data-grounded multi-turn dialogues that exhibit strong tool-calling correctness and conversational coherence. Evaluation quality is measured by applying this pipeline on seven MCP specifications spanning diverse domains and tool-suite sizes and measuring the tool-calling correctness and conversational coherence. The pipeline achieves strong quality across all domains, with complete tool coverage on small and medium specifications. Two findings emerge within this analysis: parameter schema complexity is the strongest correlate of quality variation -- tool-suite size plays a smaller, orthogonal role -- and argument value accuracy is the dominant failure mode among imperfect scenarios, a sub-dimension invisible to coarse-grained name-match metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。