工具增强推理未必更优,噪声环境下反而可能因调用开销而变差。
Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents

- 分解工具调用的三重成本:提示格式、调用协议、实际执行收益。
- 在语义干扰下,工具收益常被调用开销抵消,形成‘工具使用税’。
- 提出轻量门控机制G-STEP,缓解协议错误,但根本提升仍需模型内功。
基于大模型的智能体中,工具增强推理已成为主流方向,普遍认为其能提升推理能力与可靠性。然而我们发现,当存在语义干扰时,工具增强推理并不总优于原生思维链(CoT)。为解释这一性能差距,我们提出因子化干预框架,分离提示格式成本、工具调用协议开销及实际工具执行收益。分析显示,在语义噪声环境下,工具带来的增益往往无法弥补工具调用协议本身引入的性能下降,即‘工具使用税’。为此,我们设计G-STEP,一种轻量级推理时门控机制,以缓解协议引发的错误。尽管该方法实现部分恢复,结果表明,更显著的改进仍需增强模型内在推理与工具交互能力。
原文摘要 · Abstract (English)
Tool-augmented reasoning has become a popular direction for LLM-based agents, and it is widely assumed to improve reasoning and reliability. However, we demonstrate that this consensus does not always hold: in the presence of semantic distractors, tool-augmented reasoning does not necessarily outperform native CoT. To explain this performance gap, we propose a Factorized Intervention Framework that isolates the cost of prompt formatting, the overhead of the tool-calling protocol, and the actual gain from executing tools. Our analysis reveals a critical tradeoff: under semantic noise, the gains from tools often fail to offset the "tool-use tax", which is the performance degradation introduced by the tool-calling protocol itself. To address this, we introduce G-STEP, a lightweight inference-time gate to mitigate protocol-induced errors. While this yields partial recovery, our findings suggest that more substantial improvements still require strengthening the model's intrinsic reasoning and tool-interaction capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。