arXiv:2411.09689cs.AIcs.CL2024-11被引 3

通过内部知识扰动检测大模型幻觉,无需外部数据或微调。

Probing LLM Hallucination from Within: Perturbation-Driven Approach via Internal Knowledge

  • 基于提示中关键实体的扰动,区分三类生成结果。
  • 在四数据集四模型上超越七种现有方法,准确率领先。
  • 适合关注幻觉检测与模型可信度的研究者使用。

大模型幻觉(生成不忠实内容)严重制约其实际应用。现有检测方法多依赖外部知识、模型微调或大规模标注数据,且难以区分不同类型的幻觉。为此,我们提出幻觉探针任务,将生成文本分为一致、错位和虚构三类。基于发现:扰动提示中的关键实体会差异化影响三类文本生成,我们提出SHINE方法,无需外部知识、监督训练或模型微调。SHINE在三种现代大模型上均有效,四项基准测试中超越七种对比方法,在四数据集四模型上表现最优,验证了内部探针对精准检测的重要性。

原文摘要 · Abstract (English)

LLM hallucination, where unfaithful text is generated, presents a critical challenge for LLMs' practical applications. Current detection methods often resort to external knowledge, LLM fine-tuning, or supervised training with large hallucination-labeled datasets. Moreover, these approaches do not distinguish between different types of hallucinations, which is crucial for enhancing detection performance. To address such limitations, we introduce hallucination probing, a new task that classifies LLM-generated text into three categories: aligned, misaligned, and fabricated. Driven by our novel discovery that perturbing key entities in prompts affects LLM's generation of these three types of text differently, we propose SHINE, a novel hallucination probing method that does not require external knowledge, supervised training, or LLM fine-tuning. SHINE is effective in hallucination probing across three modern LLMs, and achieves state-of-the-art performance in hallucination detection, outperforming seven competing methods across four datasets and four LLMs, underscoring the importance of probing for accurate detection.

幻觉检测大模型内部探针SHINE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。