arXiv:2510.02837cs.AIcs.CL2025-10被引 17

提出无需参考答案的多维度评估框架,可分析工具增强型大模型推理过程。

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

  • 利用证据库累计前序步骤知识,评估推理轨迹。
  • 在小规模开源模型上仍能准确评估复杂推理过程。
  • 发现以往未报告的推理模式,适合模型开发者与评测研究者。

尽管近期的工具增强基准包含复杂任务请求,评估仍局限于答案匹配,忽略了效率、幻觉和适应性等关键轨迹特征。最直接的评估方法是将代理轨迹与真实轨迹对比,但标注所有有效真实轨迹成本过高。为此,我们提出TRACE——一种无参考的多维评估框架,用于工具增强型大语言模型。通过引入证据库累积先前步骤的知识,TRACE能有效评估代理的推理轨迹。为验证该框架,我们构建了一个新的元评估数据集,包含多样且存在缺陷的轨迹,并附有多维度性能评分。结果表明,即使使用小型开源大模型,TRACE也能准确评估复杂轨迹。此外,我们将方法应用于评估代理解决工具增强任务时的推理轨迹,揭示了此前未报告的现象及其对应洞见。

原文摘要 · Abstract (English)

Although recent tool-augmented benchmarks involve complex requests, evaluation remains limited to answer matching, neglecting critical trajectory aspects like efficiency, hallucination, and adaptivity. The most straightforward method for evaluation is to compare an agent's trajectory with the ground-truth, but annotating all valid ground-truth trajectories is prohibitively expensive. In this manner, we introduce TRACE, a reference-free framework for the multi-dimensional evaluation of tool-augmented LLMs. By incorporating an evidence bank which accumulates knowledge from preceding steps, TRACE assesses an agent's reasoning trajectory effectively. To validate our framework, we develop a new meta-evaluation dataset with diverse and flawed trajectories, each labeled with multi-faceted performance scores. Our results confirm that TRACE accurately evaluates complex trajectories even with small open-source LLMs. Furthermore, we apply our method to evaluate the trajectories that agents produce while solving tool-augmented tasks, presenting previously unreported observations and their corresponding insights.

大模型评估推理轨迹工具增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。