arXiv:2411.04847cs.CL2024-11ACL被引 12

用提示词增强大模型内部状态,提升幻觉检测跨领域泛化能力

Prompt-Guided Internal States for Hallucination Detection of Large Language Models

  • 通过提示词引导大模型内部状态中与事实性相关的结构
  • 在不同领域数据上实验,显著提升检测器泛化性能
  • 适合需要跨领域可靠幻觉检测的研究者使用

大型语言模型(LLMs)在多个领域表现出强大能力,但有时会生成逻辑连贯却事实错误的回应,即幻觉问题。现有数据驱动的监督方法依赖于模型内部状态训练检测器,但特定领域训练的检测器往往难以泛化到其他领域。本文提出一种新框架——PRISM,利用适当提示词引导模型内部状态中与文本真实性相关结构的变化,使其在不同领域的文本中更显著且一致。我们将该框架集成至现有幻觉检测方法,并在多领域数据集上进行实验。结果表明,该框架显著提升了现有检测方法的跨领域泛化能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities across a variety of tasks in different domains. However, they sometimes generate responses that are logically coherent but factually incorrect or misleading, which is known as LLM hallucinations. Data-driven supervised methods train hallucination detectors by leveraging the internal states of LLMs, but detectors trained on specific domains often struggle to generalize well to other domains. In this paper, we aim to enhance the cross-domain performance of supervised detectors with only in-domain data. We propose a novel framework, prompt-guided internal states for hallucination detection of LLMs, namely PRISM. By utilizing appropriate prompts to guide changes to the structure related to text truthfulness in LLMs' internal states, we make this structure more salient and consistent across texts from different domains. We integrated our framework with existing hallucination detection methods and conducted experiments on datasets from different domains. The experimental results indicate that our framework significantly enhances the cross-domain generalization of existing hallucination detection methods.

幻觉检测大模型提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。