构建工具学习平台ToLeaP,系统评估41个大模型的工具使用能力。
ToLeaP: Rethinking Development of Tool Learning with Large Language Models
- 搭建ToLeaP平台,实现7个模型一键评估,复现33项基准测试。
- 分析超3000个失败案例,发现大模型在自主学习、泛化和长任务解决上的四大短板。
- 提出真实场景评测、思维推理等四条新方向,推动工具学习深入发展。
工具学习使大语言模型能够有效利用外部工具,正成为提升各行业生产力的关键技术。尽管发展迅速,该领域仍存在关键挑战与机遇未被充分研究。本文通过复现33个基准测试,对41个主流大模型的工具学习能力进行评估,并构建了名为ToLeaP的工具学习平台,支持其中7个模型的一键评估。同时收集了33个潜在训练数据集中的21个,以支持未来研究。通过对超过3000个失败案例的分析,我们识别出四大核心挑战:(1) 基准测试局限导致对(2) 自主学习、(3) 泛化能力以及(4) 长时序任务求解能力的忽视。为推动未来发展,我们进一步探索四个可能方向:(1) 真实世界基准构建、(2) 兼容性感知的自主学习、(3) 通过思考实现推理学习、(4) 关键线索的识别与召回。初步实验验证了这些方向的有效性,凸显了深入研究的必要性。
原文摘要 · Abstract (English)
Tool learning, which enables large language models (LLMs) to utilize external tools effectively, has garnered increasing attention for its potential to revolutionize productivity across industries. Despite rapid development in tool learning, key challenges and opportunities remain understudied, limiting deeper insights and future advancements. In this paper, we investigate the tool learning ability of 41 prevalent LLMs by reproducing 33 benchmarks and enabling one-click evaluation for seven of them, forming a Tool Learning Platform named ToLeaP. We also collect 21 out of 33 potential training datasets to facilitate future exploration. After analyzing over 3,000 bad cases of 41 LLMs based on ToLeaP, we identify four main critical challenges: (1) benchmark limitations induce both the neglect and lack of (2) autonomous learning, (3) generalization, and (4) long-horizon task-solving capabilities of LLMs. To aid future advancements, we take a step further toward exploring potential directions, namely (1) real-world benchmark construction, (2) compatibility-aware autonomous learning, (3) rationale learning by thinking, and (4) identifying and recalling key clues. The preliminary experiments demonstrate their effectiveness, highlighting the need for further research and exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。