测试工具增强模型在信息缺失或工具不可用时的失败表现
Benchmarking Failures in Tool-Augmented Language Models
- 构建FAIL-TALMS基准,包含1749个案例与906种工具
- 多数模型无法识别缺失工具或不完整查询,仅Claude表现较好
- 实时人工协助可缓解查询不全问题,但对工具故障帮助有限
工具集成扩展了语言模型(LMs)的能力,使其超越纯文本生成。然而,当前工具增强型语言模型(TaLMs)常假设信息和工具完全可用,这在现实中并不成立。为系统研究其缺陷,我们提出FAIL-TALMS基准,涵盖两类主要失败:查询描述不充分和工具不可用。该基准包含1749个样本,涉及21类任务中的906种工具,支持单工具与多工具使用。我们评估了主流专有与开源模型,发现除Claude外,所有模型均难以识别缺失工具或信息。为进一步探索缓解策略,我们引入实时人机交互方法「Ask-and-Help」(AAH),以补充缺失信息或替换失效工具。结果显示,当查询不完整时,AAH能有效提升任务解决率;但面对复杂工具失效,其效果甚微。
原文摘要 · Abstract (English)
The integration of tools has extended the capabilities of language models (LMs) beyond vanilla text generation to versatile scenarios. However, tool-augmented language models (TaLMs) often assume 'perfect' information access and tool availability, which may not hold in the real world. To systematically study TaLMs' imperfections, we introduce the FAIL-TALMS benchmark, featuring two major failures: under-specified user queries and non-available tools. FAIL-TALMS contains 1,749 examples using 906 tools across 21 categories, including single- and multi-tool usage. We evaluate top-performing proprietary and open-source models, and find all current models except for Claude struggle to recognize missing tools or information. Further, to study possible mitigation of the failures, we enable real-time human interaction, named the Ask-and-Help (AAH) method, to provide missing information or replace non-functional tools. While AAH can help models solve tasks more correctly when queries are under-specified, it brings minimal benefit when complex tools are broken.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。