arXiv:2410.11805cs.CL2024-10中稿 · COLING 2025被引 2

构建首个评估大模型嵌套工具调用能力的数据集

NesTools: A Dataset for Evaluating Nested Tool Learning Abilities of Large Language Models

  • 自动生成多层嵌套工具调用数据,模拟真实场景
  • 22个大模型测试显示当前模型在复杂嵌套任务中表现不佳
  • 适合研究工具学习、智能代理与多步推理的学者

大型语言模型(LLMs)结合工具学习已在实际应用中取得显著成果。在工具学习过程中,模型可能以嵌套方式调用多个工具,后序工具调用可将前序响应作为输入参数。然而,现有研究对嵌套工具学习能力的探索仍不充分,因缺乏相关数据实例。为此,我们提出NesTools,填补全面评估嵌套工具学习能力的空白。NesTools采用创新的自动数据生成方法,构建大规模具有不同嵌套结构的工具调用数据。经人工审核与优化,数据质量高且贴近真实场景。因此,NesTools可作为新基准,用于评估LLMs的嵌套工具学习能力。我们在22个LLMs上开展广泛实验,并基于NesTools进行深入分析,结果表明当前大模型在复杂嵌套工具学习任务中仍存在明显短板。

原文摘要 · Abstract (English)

Large language models (LLMs) combined with tool learning have gained impressive results in real-world applications. During tool learning, LLMs may call multiple tools in nested orders, where the latter tool call may take the former response as its input parameters. However, current research on the nested tool learning capabilities is still under-explored, since the existing benchmarks lack relevant data instances. To address this problem, we introduce NesTools to bridge the current gap in comprehensive nested tool learning evaluations. NesTools comprises a novel automatic data generation method to construct large-scale nested tool calls with different nesting structures. With manual review and refinement, the dataset is in high quality and closely aligned with real-world scenarios. Therefore, NesTools can serve as a new benchmark to evaluate the nested tool learning abilities of LLMs. We conduct extensive experiments on 22 LLMs, and provide in-depth analyses with NesTools, which shows that current LLMs still suffer from the complex nested tool learning task.

工具学习大模型评测嵌套调用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。