TIDE-Bench提升工具增强型推理评估质量,覆盖多类复杂任务与诊断指标。
TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning

- 设计数学推理与动态交互等四类任务,涵盖工具调用与多工具协作能力
- 综合评估答案准确率、过程可靠性、工具使用效率和推理成本,覆盖异构任务场景
- 筛选高区分度样本,降低评估成本,适合研究工具集成推理的学者使用
工具集成推理(TIR)通过引入外部计算、检索与执行能力,成为提升大语言模型性能的重要范式。然而,当前领域仍缺乏高质量、统一的评估基准,现有评价在数据集质量、任务多样性、诊断全面性与评估效率方面存在局限。本文提出TIDE-Bench,一个全面且高效的TIR评估基准,具备三大优势:首先,涵盖广泛使用的数学推理与知识密集型问答任务,以及新设计的工具引导实验设计任务和动态交互任务,以检验模型在复杂工具调用与多工具协同中的表现;其次,采用任务感知的综合性评估协议,联合衡量最终答案质量、过程可靠性、工具使用效率与推理开销;第三,通过过滤低区分度样本构建高质量评估集,在显著降低评估成本的同时聚焦更具挑战性的样本。对多个基础模型与TIR方法的大量实验揭示了工具定位方面的持续瓶颈,为未来研究提供重要洞见。
原文摘要 · Abstract (English)
Tool-integrated reasoning has emerged as a promising paradigm for enhancing large language models with external computation, retrieval, and execution capabilities. However, the field still lacks a high-quality and unified evaluation benchmark, and existing TIR evaluations remain limited in dataset quality, task diversity, diagnostic comprehensiveness, and evaluation efficiency. In this work, we introduce TIDE-Bench, a holistic and efficient benchmark for evaluating TIR methods, featuring three key advantages. First, it provides diverse task settings, combining widely used mathematical reasoning and knowledge-intensive QA tasks with two newly designed tasks, namely the tool-grounded experimental design task and the dynamic interactive task, to probe models' abilities in complex tool invocation and multi-tool coordination. Second, TIDE-Bench adopts a comprehensive yet task-aware evaluation protocol, jointly measuring final answer quality, process reliability, tool-use efficiency, and inference cost across heterogeneous task settings. Third, TIDE-Bench constructs high-quality and discriminative evaluation sets by filtering low-discrimination instances from existing datasets, substantially reducing evaluation cost while focusing on more challenging samples. Extensive experiments on multiple foundation models and TIR methods reveal persistent bottlenecks in tool grounding, offering insights for future TIR research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。