用结构化模板提升大模型函数调用准确率与可解释性
Improving Large Language Models Function Calling and Interpretability via Guided-Structured Templates
- 设计课程式结构化模板引导模型分步生成函数调用
- 在多个模型上减少工具使用错误,相对基线提升3-12%
- 增强AI助手的可解释性,适合需要可靠工具调用的应用
大语言模型虽具备强大推理与工具使用能力,但在真实场景中常因参数设置错误、工具选择不当或用户意图误解而失败。这些问题多源于对用户目标理解不全及对工具文档认知不足。尽管思维链(CoT)提示在一般推理中有效,但我们的分析表明,自由格式的CoT在结构化函数调用任务中效果有限,甚至可能适得其反。为此,我们提出一种受课程启发的框架,利用结构化推理模板引导模型更严谨地生成函数调用。实验显示,该方法显著降低工具使用错误,在多种模型系列和方法上实现3%-12%的相对改进。同时,框架提升了工具调用的鲁棒性、可解释性与透明度,推动更可靠的现实应用级AI助手发展。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated strong reasoning and tool-use capabilities, yet they often fail in real-world tool-interactions due to incorrect parameterization, poor tool selection, or misinterpretation of user intent. These issues often stem from an incomplete understanding of user goals and inadequate comprehension of tool documentation. While Chain-of-Thought (CoT) prompting has proven effective for enhancing reasoning in general contexts, our analysis reveals that free-form CoT is insufficient and sometimes counterproductive for structured function-calling tasks. To address this, we introduce a curriculum-inspired framework that leverages structured reasoning templates to guide LLMs through more deliberate step-by-step instructions for generating function callings. Experimental results show that our method reduces tool-use errors, achieving 3-12% relative improvements over strong baselines across diverse model series and approaches. Moreover, our framework enhances the robustness, interpretability, and transparency of tool-using agents, advancing the development of more reliable AI assistants for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。