arXiv:2506.13977cs.SEcs.CL2025-06EMNLP被引 19

评测大模型在工具调用出错时的自我纠错能力,构建新基准提升评估可信度。

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

  • 设计进化式数据构造策略生成多样错误场景。
  • 在多个基准上验证模型纠错能力,发现主流模型表现有限。
  • 适合研究大模型工具调用与自我反思的学者使用。

大型语言模型(LLMs)利用外部工具的能力使其能处理日益多样的任务。然而,随着任务复杂性和长时程需求增加,工具调用过程可能出现多种意外错误。如何有效识别、诊断并恢复这些错误,已成为推动工具学习的关键方向。本文首先在多个前沿工具评估基准上系统分析了函数调用过程中的错误类型。基于此,提出 CRITICTOOL,一个专为工具学习设计的综合性批判评估基准。该基准采用新颖的演化策略构建,涵盖不同复杂度的工具使用错误,更贴近真实应用场景。我们在 CRITICTOOL 上开展广泛实验,验证了所提基准构建策略的泛化性与有效性,并深入分析了各类大模型在工具调用中的反思能力,为大模型工具学习研究提供了新视角。代码已公开于 https://github.com/Shellorley0513/CriticTool。

原文摘要 · Abstract (English)

The ability of large language models (LLMs) to utilize external tools has enabled them to tackle an increasingly diverse range of tasks. However, as the tasks become more complex and long-horizon, the intricate tool utilization process may trigger various unexpected errors. Therefore, how to effectively handle such errors, including identifying, diagnosing, and recovering from them, has emerged as a key research direction for advancing tool learning. In this work, we first extensively analyze the types of errors encountered during the function-calling process on several competitive tool evaluation benchmarks. Based on it, we introduce CRITICTOOL, a comprehensive critique evaluation benchmark specialized for tool learning. Building upon a novel evolutionary strategy for dataset construction, CRITICTOOL holds diverse tool-use errors with varying complexities, which better reflects real-world scenarios. We conduct extensive experiments on CRITICTOOL, and validate the generalization and effectiveness of our constructed benchmark strategy. We also provide an in-depth analysis of the tool reflection ability on various LLMs, offering a new perspective on the field of tool learning in LLMs. The code is available at \href{https://github.com/Shellorley0513/CriticTool}{https://github.com/Shellorley0513/CriticTool}.

大模型工具调用自省能力评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。