arXiv:2505.00557cs.CLcs.AI2025-05被引 2

通过误导性提示量化大模型幻觉,揭示其不靠谱的根源。

Triggering Hallucinations in LLMs: A Quantitative Study of Prompt-Induced Hallucination in Large Language Models

  • 设计误导性提示融合无关概念,诱发幻觉
  • 不同模型幻觉程度差异明显,推理型模型更易受影响
  • 可复现测试框架,适合安全与可靠性研究者

大语言模型(LLMs)的幻觉问题在医疗、法律等真实场景中日益突出,尽管经过对齐和指令微调,模型仍可能生成流畅但错误的内容。本文提出一种基于提示的系统化方法:通过合成融合语义相距甚远的概念(如元素周期表与塔罗占卜)的幻觉诱导提示(HIP),并用幻觉量化提示(HQP)评估输出的合理性、置信度与连贯性。在多个大模型上的对照实验表明,使用HIP的输出比对照组更不连贯且幻觉更严重,且不同模型表现各异,推理类模型表现出独特模式。该框架为研究幻觉脆弱性提供了可复现的测试平台,有助于开发能自我检测和调节概念不稳定的更安全模型。

原文摘要 · Abstract (English)

Hallucinations in large language models (LLMs) present a growing challenge across real-world applications, from healthcare to law, where factual reliability is essential. Despite advances in alignment and instruction tuning, LLMs can still generate outputs that are fluent yet fundamentally untrue. Understanding the cognitive dynamics that underlie these hallucinations remains an open problem. In this study, we propose a prompt-based framework to systematically trigger and quantify hallucination: a Hallucination-Inducing Prompt (HIP), which synthetically fuses semantically distant concepts (e.g., periodic table of elements and tarot divination) in a misleading way, and a Hallucination Quantifying Prompt (HQP), which scores the plausibility, confidence, and coherence of the output. Controlled experiments across multiple LLMs revealed that HIPs consistently produced less coherent and more hallucinated responses than their null-fusion controls. These effects varied across models, with reasoning-oriented LLMs showing distinct profiles from general-purpose ones. Our framework provides a reproducible testbed for studying hallucination vulnerability, and opens the door to developing safer, more introspective LLMs that can detect and self-regulate the onset of conceptual instability.

幻觉检测提示工程大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。