arXiv:2503.01763cs.CLcs.AI2025-03ACL被引 50

评测大模型工具检索能力,发现现有模型表现差,影响任务成功率。

Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models

  • 构建包含7600个任务、4.3万工具的异构检索基准
  • 主流模型在该任务上表现不佳,任务通过率下降
  • 提供20万+训练数据,显著提升工具检索效果

工具学习旨在为大语言模型(LLMs)注入多样化工具,使其能作为智能体解决实际任务。由于工具使用场景下上下文长度受限,采用信息检索(IR)模型从大规模工具集中筛选有用工具成为关键第一步。然而,当前对IR模型在工具检索任务中的性能缺乏充分评估,多数工具使用基准通过人工预标注少量相关工具简化此步骤,与真实场景严重脱节。本文提出ToolRet,一个包含7600个多样化检索任务和4.3万工具的异构工具检索基准,涵盖现有数据集收集的工具。我们在ToolRet上对六类模型进行评测,结果令人意外:即使在传统IR基准中表现优异的模型,在ToolRet上也表现不佳。低质量的工具检索导致工具使用型大模型的任务通过率下降。为进一步改进,我们构建了一个超大规模训练数据集,包含超过20万条样本,显著提升了IR模型的工具检索能力。

原文摘要 · Abstract (English)

Tool learning aims to augment large language models (LLMs) with diverse tools, enabling them to act as agents for solving practical tasks. Due to the limited context length of tool-using LLMs, adopting information retrieval (IR) models to select useful tools from large toolsets is a critical initial step. However, the performance of IR models in tool retrieval tasks remains underexplored and unclear. Most tool-use benchmarks simplify this step by manually pre-annotating a small set of relevant tools for each task, which is far from the real-world scenarios. In this paper, we propose ToolRet, a heterogeneous tool retrieval benchmark comprising 7.6k diverse retrieval tasks, and a corpus of 43k tools, collected from existing datasets. We benchmark six types of models on ToolRet. Surprisingly, even the models with strong performance in conventional IR benchmarks, exhibit poor performance on ToolRet. This low retrieval quality degrades the task pass rate of tool-use LLMs. As a further step, we contribute a large-scale training dataset with over 200k instances, which substantially optimizes the tool retrieval ability of IR models.

工具学习信息检索大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。