真实用户行为让大模型工具使用难上加难,新基准揭示模型能力严重不足。
Benchmarking LLM Tool-Use in the Wild

- 基于真实对话设计评测框架,捕捉复杂多步工具调用
- 57个模型平均准确率不足15%,暴露代理能力短板
- 适合研究大模型交互鲁棒性与真实场景应用的学者
通过大语言模型实现用户需求的多轮、多步骤工具调用,往往并非直截了当。真实用户交互本质上是复杂、混乱且灵活的。我们识别出三大用户行为挑战:需要高效编排工具调用拓扑的组合任务;意图分散在多轮对话中需上下文推断;指令切换混合任务查询、澄清与闲聊,迫使模型实时调整策略。现有基准忽略这些行为,导致模型工具使用性能提升看似进步实则虚假。为此,我们提出基于真实用户行为模式的WildToolBench基准。对57个LLM的全面评估显示,无一模型准确率超过15%,表明大模型代理能力存在显著差距。控制实验与深入分析进一步表明,真正挑战不在于人为设计的复杂任务,而在于用户行为的‘野生’特性,强调需重新思考大模型、用户与工具间的互动机制。
原文摘要 · Abstract (English)
Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inherently wild, being intricate, messy, and flexible. We identify three key challenges from user behaviour: compositional tasks that demand efficient orchestration of tool-call topologies, implicit intent spread across dialogue turns that require contextual inference, and instruction transition, which mixes task queries, clarifications, and casual conversation, forcing LLMs to adjust their policies on the fly. Existing benchmarks overlook these behaviors, making the apparent progress of LLMs on tool-use spurious. To address this, we introduce WildToolBench, an LLM tool-use benchmark grounded in real-world user behavior patterns. Comprehensive evaluations of 57 LLMs reveal that no model achieves an accuracy of more than 15%, indicating a substantial gap in the robustness of LLMs' agentic ability. Controlled experiments and in-depth analyses further indicate that the real challenge for LLM tool-use lies not in artificially complex tasks, but in the wild nature of user behavior, emphasizing the need to reconsider the interactions among LLMs, users, and tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。