评测大模型工具菜单过滤策略,提升智能体可靠性和效率。
ToolMenuBench: Benchmarking Tool-Menu Filtering Strategies for Reliable and Efficient LLM Agents

- 设计多维度工具菜单筛选测试框架,覆盖大小、干扰项等变量。
- 最优策略使任务成功率从32.1%提至85.7%,耗能降低98%。
- 适合研究智能体界面设计与安全风险控制的开发者与研究人员。
工具增强型大语言模型智能体在大型工具库上运行时,现有评估常关注是否调用正确工具,而忽视可见工具菜单对可靠性、效率和安全风险的影响。我们提出ToolMenuBench,一个用于评估多步骤大模型智能体中工具菜单构建的基准测试。该基准涵盖工具菜单规模、干扰类型、状态依赖任务结构及风险暴露等维度,并报告过滤级别与下游智能体指标,包括可见工具数、高风险工具暴露、任务成功率、错误调用、提前动作和令牌使用量。在七种模型后端、三种菜单规模、六种过滤方法及七个评估设置的受控实验中,因果最小工具过滤(CMTF)将任务成功率从全工具暴露下的32.1%提升至85.7%,平均令牌使用量减少约98%。该方法在减少可见工具数、错误调用、提前动作及高风险暴露方面均优于无过滤、词法过滤、状态感知过滤以及更广泛因果路径基线。ToolMenuBench提供可复用的评估框架,以解决智能体-界面问题:何时、何地、在何种成本或风险约束下应显示哪些工具。
原文摘要 · Abstract (English)
Tool-augmented large language model agents increasingly operate over large tool libraries, but existing evaluations often focus on whether a model can call a tool correctly rather than how the visible tool menu shapes reliability, efficiency, and safety-relevant risk exposure. We introduce ToolMenuBench, a benchmark for evaluating tool-menu construction in multi-step LLM agents. ToolMenuBench varies tool-menu size, distractor type, state-dependent task structure, and risk exposure, and reports both filter-level and downstream agent metrics, including visible-tool count, risky-tool exposure, task success, wrong-tool calls, premature actions, and token usage. In a controlled evaluation across seven model backends, three tool-menu sizes, six filtering methods, and seven evaluation settings, CMTF improves task success from 32.1% under all-tools exposure to 85.7%, while reducing average token usage by roughly 98%. Causal minimal tool filtering achieves the strongest overall tradeoff, reducing visible tools, wrong-tool calls, premature actions, and risky-tool exposure relative to unfiltered exposure, lexical filtering, state-aware filtering, and broader causal-path baselines. ToolMenuBench provides a reusable evaluation framework for studying the agent-interface problem: which tools should be visible, when they should be visible, and under what cost or risk constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。