arXiv:2508.15754cs.CLcs.AI2025-08被引 2

TIR让大模型更准更快地推理,尤其在需要工具的任务上表现突出。

Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis

  • 引入九类任务的ReasonZoo基准,评估工具融合推理效果。
  • 使用新指标PAC和AUC-PCC发现TIR减少冗余思考,提升效率。
  • 在数学与非数学任务中均显著优于无工具模型,适用广泛。

大语言模型(LLMs)通过链式思维(CoT)推理取得进展,但在精确计算任务中仍表现不足。工具集成推理(TIR)通过引入外部工具改善这一问题,但其在提升模型推理能力方面的泛化性尚不明确,且对模型思维方式的影响也未被深入研究。为此,我们构建了包含九类多样化推理任务的ReasonZoo基准,用于评估TIR在多领域中的有效性。同时提出两个新指标:性能感知成本(PAC)与性能-成本曲线下的面积(AUC-PCC),以衡量推理效率。实证结果表明,启用TIR的模型在数学与非数学任务中均持续优于非TIR模型。此外,TIR显著提升推理效率,表现为更高的PAC与AUC-PCC值,说明减少了过度思考,使推理过程更精简。这些发现证实了TIR具有领域通用优势,有望推动大模型在复杂推理任务中的能力发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have made significant strides in reasoning tasks through methods like chain-of-thought (CoT) reasoning. However, they often fall short in tasks requiring precise computations. Tool-Integrated Reasoning (TIR) has emerged as a solution by incorporating external tools into the reasoning process. Nevertheless, the generalization of TIR in improving the reasoning ability of LLM is still unclear. Additionally, whether TIR has improved the model's reasoning behavior and helped the model think remains to be studied. We introduce ReasonZoo, a comprehensive benchmark encompassing nine diverse reasoning categories, to evaluate the effectiveness of TIR across various domains. Additionally, we propose two novel metrics, Performance-Aware Cost (PAC) and Area Under the Performance-Cost Curve (AUC-PCC), to assess reasoning efficiency. Our empirical evaluation demonstrates that TIR-enabled models consistently outperform their non-TIR counterparts in both mathematical and non-mathematical tasks. Furthermore, TIR enhances reasoning efficiency, as evidenced by improved PAC and AUC-PCC, indicating reduced overthinking and more streamlined reasoning. These findings underscore the domain-general benefits of TIR and its potential to advance LLM capabilities in complex reasoning tasks.

工具推理大模型效率评估推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。