arXiv:2507.15296cs.SEcs.AI2025-07EMNLP被引 16

分析大模型工具链中参数填充失败问题,揭示根源并提出改进方案。

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems

  • 构建工具调用链的参数失败分类体系,归纳出五类典型错误。
  • 实验发现参数名幻觉主要源于模型自身局限,输入源差异影响其他错误类型。
  • 建议统一返回格式、增强错误反馈、保障参数一致性以提升可靠性。

工具代理范式拓展了大语言模型的能力边界,使其能完成更复杂的任务。然而,执行过程中参数填充失败限制了该范式的有效性。本文首先构建了一个参数失败分类体系,基于主流工具代理的调用链归纳出五类失败模式。通过在输入上应用15种扰动方法,探究三类不同输入源与失败类别之间的关联。实验结果表明,参数名幻觉主要源自模型固有缺陷,而输入源差异则主导其他失败模式。为提升工具-代理交互的可靠性和有效性,本文提出三项改进建议:标准化工具返回格式、优化错误反馈机制、确保参数一致性。

原文摘要 · Abstract (English)

The emergence of the tool agent paradigm has broadened the capability boundaries of the Large Language Model (LLM), enabling it to complete more complex tasks. However, the effectiveness of this paradigm is limited due to the issue of parameter failure during its execution. To explore this phenomenon and propose corresponding suggestions, we first construct a parameter failure taxonomy in this paper. We derive five failure categories from the invocation chain of a mainstream tool agent. Then, we explore the correlation between three different input sources and failure categories by applying 15 input perturbation methods to the input. Experimental results show that parameter name hallucination failure primarily stems from inherent LLM limitations, while issues with input sources mainly cause other failure patterns. To improve the reliability and effectiveness of tool-agent interactions, we propose corresponding improvement suggestions, including standardizing tool return formats, improving error feedback mechanisms, and ensuring parameter consistency.

大模型工具链参数失败

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。