arXiv:2506.00042cs.CL2025-06ACL被引 7

用分层检查表提升大模型调用工具的准确性

Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists

  • 设计全局与局部双重错误检查表,系统诊断工具调用问题
  • 在五个数据集上参数填充准确率显著优于基线方法
  • 适合需要高可靠工具调用的智能助手、自动化系统

大型语言模型在自然语言处理中取得显著进展,尤其体现在集成外部工具和API方面。然而,其效果常因工具调用时参数填写错误而受阻。本文提出分层工具错误检查表(HiTEC)框架,无需依赖大量真实交互即可系统诊断并缓解工具调用错误。该框架采用两级策略:全局错误检查表识别跨工具常见问题,局部检查表聚焦特定工具与上下文失败。基于此结构,提出两种部署方式:HiTEC-ICL将全局检查表嵌入初始提示,通过两轮对话动态优化参数处理;HiTEC-KTO生成高质量负例,通过基于偏好优化驱动微调。在五个公开数据集上的实验表明,本框架在参数填充准确率和工具调用成功率方面均显著优于基线方法。

原文摘要 · Abstract (English)

Large language models (LLMs) have significantly advanced natural language processing, particularly through the integration of external tools and APIs. However, their effectiveness is frequently hampered by parameter mis-filling during tool calling. In this paper, we propose the Hierarchical Tool Error Checklist (HiTEC) framework to systematically diagnose and mitigate tool-calling errors without relying on extensive real-world interactions. HiTEC introduces a two-tiered approach: a global error checklist that identifies common, cross-tool issues, and a local error checklist that targets tool-specific and contextual failures. Building on this structure, we propose two deployments: HiTEC-In Context Learning (HiTEC-ICL) and HiTEC-Kahneman-Tversky Optimization (HiTEC-KTO). HiTEC-ICL embeds the global checklist in the initial prompts and leverages a two-round conversational interaction to dynamically refine parameter handling, while HiTEC-KTO generates high-quality negative examples to drive fine-tuning via preference-based optimization. Extensive experiments across five public datasets demonstrate that our framework significantly improves parameter-filling accuracy and tool-calling success rates compared to baseline methods.

大模型工具调用错误检测提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。