arXiv:2505.16986cs.CLcs.AI2025-05NeurIPS被引 14

构建多轮工具调用对话数据集,提升智能体规划能力

T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning

  • 设计支持跨工具依赖的多轮对话数据集
  • 涵盖9个领域,支持动态重规划与缓存机制
  • 适合作为大模型工具使用能力评估基准

大型语言模型在作为智能体解决复杂问题方面已展现强大能力。然而,在涉及API或工具调用依赖关系的多轮对话场景中,有效规划仍是重大挑战。为此,我们提出T1,一个面向工具增强、跨领域、多轮对话的数据集,专门用于捕捉和管理不同领域中的跨工具依赖。T1支持对九个不同领域(4个单领域,5个跨领域)的严格评估,集成短时与长时记忆缓存机制,并支持动态重规划——如判断是否重新计算或复用缓存结果。除推动工具使用与规划研究外,T1还可作为开放权重与专有大模型性能评估的基准。我们基于T1-Agent展示的结果表明,该模型可在复杂工具依赖场景中实现有效规划与推理。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated impressive capabilities as intelligent agents capable of solving complex problems. However, effective planning in scenarios involving dependencies between API or tool calls-particularly in multi-turn conversations-remains a significant challenge. To address this, we introduce T1, a tool-augmented, multi-domain, multi-turn conversational dataset specifically designed to capture and manage inter-tool dependencies across diverse domains. T1 enables rigorous evaluation of agents' ability to coordinate tool use across nine distinct domains (4 single domain and 5 multi-domain) with the help of an integrated caching mechanism for both short- and long-term memory, while supporting dynamic replanning-such as deciding whether to recompute or reuse cached results. Beyond facilitating research on tool use and planning, T1 also serves as a benchmark for evaluating the performance of open-weight and proprietary large language models. We present results powered by T1-Agent, highlighting their ability to plan and reason in complex, tool-dependent scenarios.

工具调用多轮对话智能体规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。