arXiv:2412.04141cs.CL2024-12ICML被引 49

通过允许模型拒绝使用工具,显著降低大模型工具幻觉问题。

Reducing Tool Hallucination via Reliability Alignment

  • 引入可延迟或调整工具使用的新型动作空间
  • 在测试中使工具幻觉率下降,任务成功率提升
  • 适合需要高可靠性的自动化系统开发者

大型语言模型已从单纯的语言生成扩展到与外部工具交互,推动自动化和真实场景应用。然而,工具幻觉——即模型选择不恰当工具或误用工具——带来严重挑战,导致任务错误执行、计算成本上升及系统可靠性下降。为此,本文定义并分类两种主要幻觉类型:工具选择幻觉与工具使用幻觉。为评估和缓解该问题,提出RelyToolBench,集成专用测试用例和新指标,用于衡量幻觉感知下的任务成功与效率。进一步提出Relign框架,将工具使用动作空间扩展至包含犹豫、求助与动态调整等非确定性行为,使模型能主动推迟或修正工具调用。大量实验表明,Relign有效减少工具幻觉,提升任务可靠性与交互效率。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have expanded their capabilities beyond language generation to interact with external tools, enabling automation and real-world applications. However, tool hallucinations, where models either select inappropriate tools or misuse them, pose significant challenges, leading to erroneous task execution, increased computational costs, and reduced system reliability. To systematically address this issue, we define and categorize tool hallucinations into two main types, tool selection hallucination and tool usage hallucination. To evaluate and mitigate these issues, we introduce RelyToolBench, which integrates specialized test cases and novel metrics to assess hallucination-aware task success and efficiency. Finally, we propose Relign, a reliability alignment framework that expands the tool-use action space to include indecisive actions, allowing LLMs to defer tool use, seek clarification, or adjust tool selection dynamically. Through extensive experiments, we demonstrate that Relign significantly reduces tool hallucinations, improves task reliability, and enhances the efficiency of LLM tool interactions.

大模型工具调用幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。