arXiv:2607.14108cs.CLcs.AI2026-07

提出工具效率新度量,量化大模型调用工具的有用性。

Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility

论文配图:Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
图 1 · 摘自论文原文
  • 用边际工具效用判断每次调用是否必要
  • 通过大模型判别器识别可删减的冗余工具
  • 适合优化工具套件、提升代理运行效率

本文提出工具效率这一新量化指标,用于评估大模型智能体轨迹中有效工具调用的比例。为确保该指标可定义,引入边际工具效用,对每次工具调用进行量化评估,判断其是否真正有用或可安全移除而不影响准确率,从而提升效率。本文采用大模型作为裁判(LLM-as-a-Judge)来确定轨迹中每一步工具调用的边际工具效用符号。现有研究多通过准确率间接衡量工具使用效率,而本工作首次在事后分析中直接测量效率。本研究旨在推动大模型评估前沿,为未来基准设计与智能体工程(特别是构建精简工具集)提供可优化的独立于准确率的新指标。

原文摘要 · Abstract (English)

This paper introduces tool efficiency, a new quantitative metric to evaluate the rate of useful tool calls in an LLM agent trajectory. To ensure that tool efficiency is well-defined, we also introduce marginal tool utility, a new quantitative metric defined per tool call indicating whether a tool is useful or whether it can be safely removed from the tool suite without affecting accuracy while increasing tool efficiency; in this paper, we determine the sign of marginal tool utility for each tool call in a trajectory using LLM-as-a-Judge. While much prior work has been done to develop techniques that improve tool use by LLMs and design evaluation methods measuring efficiency indirectly using accuracy as a proxy, our work is centered on measuring efficiency directly via the quantitative metric proposed in this paper in post hoc trajectory analyses. It is our intention that this work contributes to the frontier of LLM evaluation research as a springboard for future benchmark designs and agent harness engineering (specifically with regards to creating lean tool suites) that optimize for metrics that complement but are distinct from accuracy.

大模型评估工具效率智能体优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。