arXiv:2505.13328cs.CL2025-05ACL被引 16

构建多轮有状态工具交互数据集,揭示大模型长期工具使用能力不足。

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges

  • 设计多轮对话中完整的工具生命周期评估框架
  • 13个模型在长程任务中表现不佳,最高仅62%正确率
  • 适合研究智能体长期记忆与工具调用的学者参考

现有评估基准大多聚焦于无状态、单轮的工具使用或部分评价(如单轮工具选择),忽视了多轮应用中交互的固有状态特性。为弥补这一空白,我们提出 exttt{DialogTool}——一个涵盖六项关键任务、三个阶段的多轮有状态工具交互数据集:1)工具创建;2)工具使用:包括工具感知、选择与执行;3)角色一致响应:响应生成与角色扮演。此外,我们构建了 exttt{VirtualMobile}——一个具身化的虚拟移动环境,用于模拟 API 调用并评估所生成 API 的鲁棒性(本文中工具与 API 可互换使用)。利用这些资源,我们对13个开源与闭源大模型进行了全面评估,并在各阶段提供详细分析,结果表明当前最先进的大模型在长周期工具使用中仍表现不佳。

原文摘要 · Abstract (English)

Existing benchmarks that assess Language Models (LMs) as Language Agents (LAs) for tool use primarily focus on stateless, single-turn interactions or partial evaluations, such as tool selection in a single turn, overlooking the inherent stateful nature of interactions in multi-turn applications. To fulfill this gap, we propose \texttt{DialogTool}, a multi-turn dialogue dataset with stateful tool interactions considering the whole life cycle of tool use, across six key tasks in three stages: 1) \textit{tool creation}; 2) \textit{tool utilization}: tool awareness, tool selection, tool execution; and 3) \textit{role-consistent response}: response generation and role play. Furthermore, we build \texttt{VirtualMobile} -- an embodied virtual mobile evaluation environment to simulate API calls and assess the robustness of the created APIs\footnote{We will use tools and APIs alternatively, there are no significant differences between them in this paper.}. Taking advantage of these artifacts, we conduct comprehensive evaluation on 13 distinct open- and closed-source LLMs and provide detailed analysis at each stage, revealing that the existing state-of-the-art LLMs still cannot perform well to use tools over long horizons.

语言模型工具使用多轮对话评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。