arXiv:2410.19385cs.CLcs.AI2024-10被引 14

研究提示词与外部工具如何影响大模型幻觉率,发现简单提示更有效,工具调用反而增加错误。

Investigating the Role of Prompting and External Tools in Hallucination Rates of Large Language Models

  • 对比多种提示策略,测试其在不同任务中的幻觉抑制效果。
  • 结果显示简单提示常优于复杂方法,且工具调用使幻觉率显著上升。
  • 适合关注大模型可靠性与提示工程优化的研究者和开发者。

大型语言模型(LLMs)在大量人类可读文本上训练,具备通用语言理解与生成能力,在自然语言处理任务中表现优异。然而,它们常产生不准确内容,即所谓‘幻觉’。提示工程作为设计指令以引导模型完成特定任务的方法,已成为缓解幻觉的关键手段。本文对多种提示策略与框架进行系统性实证评估,涵盖广泛基准数据集,比较各方法的准确性与幻觉率。同时,研究了工具调用代理(通过外部工具增强能力的LLM)在相同基准上的幻觉表现。结果表明,最优提示策略取决于问题类型,且简单方法往往优于复杂方案;此外,引入外部工具会显著提高幻觉率,反映其带来的额外复杂性风险。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are powerful computational models trained on extensive corpora of human-readable text, enabling them to perform general-purpose language understanding and generation. LLMs have garnered significant attention in both industry and academia due to their exceptional performance across various natural language processing (NLP) tasks. Despite these successes, LLMs often produce inaccuracies, commonly referred to as hallucinations. Prompt engineering, the process of designing and formulating instructions for LLMs to perform specific tasks, has emerged as a key approach to mitigating hallucinations. This paper provides a comprehensive empirical evaluation of different prompting strategies and frameworks aimed at reducing hallucinations in LLMs. Various prompting techniques are applied to a broad set of benchmark datasets to assess the accuracy and hallucination rate of each method. Additionally, the paper investigates the influence of tool-calling agents (LLMs augmented with external tools to enhance their capabilities beyond language generation) on hallucination rates in the same benchmarks. The findings demonstrate that the optimal prompting technique depends on the type of problem, and that simpler techniques often outperform more complex methods in reducing hallucinations. Furthermore, it is shown that LLM agents can exhibit significantly higher hallucination rates due to the added complexity of external tool usage.

大模型幻觉抑制提示工程工具调用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。