arXiv:2412.12151cs.LGcs.AI2024-12EMNLP被引 16

提出SMARTCAL方法,让大模型更清醒地使用工具

SMARTCAL: An Approach to Self-Aware Tool-Use Evaluation and Calibration

  • 设计自感知工具使用评估与校准框架
  • 提升问答准确率8.6%,降低校准误差21.6%
  • 适合关注模型可信度与工具调用安全的研究者

大语言模型的工具使用能力对工业应用影响深远,但其在适切使用工具时的自我控制与校准能力仍缺乏研究。本文在三个数据集上,针对两类主流工具使用框架,对一系列先进大模型进行了分析。研究发现,大模型普遍存在工具滥用行为,表现为高自信下过度使用工具,且该问题在不同能力模型中均存在。为此,我们提出新型方法SMARTCAL以缓解上述问题。实验结果表明,相比基线模型,该方法平均提升问答性能8.6%,并使期望校准误差(ECE)下降21.6%。

原文摘要 · Abstract (English)

The tool-use ability of Large Language Models (LLMs) has a profound impact on a wide range of industrial applications. However, LLMs' self-control and calibration capability in appropriately using tools remains understudied. The problem is consequential as it raises potential risks of degraded performance and poses a threat to the trustworthiness of the models. In this paper, we conduct a study on a family of state-of-the-art LLMs on three datasets with two mainstream tool-use frameworks. Our study reveals the tool-abuse behavior of LLMs, a tendency for models to misuse tools with overconfidence. We also find that this is a common issue regardless of model capability. Accordingly, we propose a novel approach, \textit{SMARTCAL}, to mitigate the observed issues, and our results show an average of 8.6 percent increase in the QA performance and a 21.6 percent decrease in Expected Calibration Error (ECE) compared to baseline models.

大模型工具使用校准评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。