新基准ToolScan揭示大模型工具使用中的七类错误模式。
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
- 构建多环境查询数据集,识别工具使用中的七类错误
- 主流大模型在任务中普遍存在这些错误模式
- 帮助研究者针对性改进模型可靠性,适合系统开发者
评估大语言模型(LLMs)是构建高性能复合人工智能系统的关键。由于LLM的输出会传递到下游步骤,识别其错误对系统性能至关重要。在AI系统中,工具使用是LLM的常见任务。尽管已有多个基准环境用于评估此任务,但通常仅提供成功率,缺乏对失败案例的分析。为此,我们提出TOOLSCAN,一个用于识别LLM在工具使用任务中输出错误模式的新基准。该基准数据集包含来自多种环境的查询,可用于检测七种新识别的错误模式。利用TOOLSCAN,我们发现即使最领先的LLMs在其输出中也表现出这些错误模式。研究人员可借助TOOLSCAN的洞察,指导误差缓解策略的制定。
原文摘要 · Abstract (English)
Evaluating Large Language Models (LLMs) is one of the most critical aspects of building a performant compound AI system. Since the output from LLMs propagate to downstream steps, identifying LLM errors is crucial to system performance. A common task for LLMs in AI systems is tool use. While there are several benchmark environments for evaluating LLMs on this task, they typically only give a success rate without any explanation of the failure cases. To solve this problem, we introduce TOOLSCAN, a new benchmark to identify error patterns in LLM output on tool-use tasks. Our benchmark data set comprises of queries from diverse environments that can be used to test for the presence of seven newly characterized error patterns. Using TOOLSCAN, we show that even the most prominent LLMs exhibit these error patterns in their outputs. Researchers can use these insights from TOOLSCAN to guide their error mitigation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。