区分内在与外在幻觉,用注意力机制提升检测精度。
The Map of Misbelief: Tracing Intrinsic and Extrinsic Hallucinations Through Attention Patterns
- 基于注意力模式聚合,区分内外部幻觉类型。
- 相比采样方法,该策略对内在幻觉检测准确率更高。
- 适合关注模型可信度与可解释性的研究者。
大型语言模型在安全关键领域应用日益广泛,但仍易产生幻觉。现有幻觉检测方法多依赖计算开销大的采样策略,且常忽略幻觉类型的差异。本文提出一个系统性评估框架,区分外在与内在幻觉,并在一组精心设计的基准上评估检测性能。我们引入一种基于注意力的不确定性量化算法,提出新型注意力聚合策略,同时提升可解释性与检测效果。实验表明,语义熵等采样方法对检测外在幻觉有效,但对内在幻觉普遍失效;而我们的注意力聚合方法在内在幻觉检测上表现更优。这些发现为匹配检测策略与幻觉本质提供了新方向,凸显注意力是量化模型不确定性的丰富信号。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed in safety-critical domains, yet remain susceptible to hallucinations. While prior works have proposed confidence representation methods for hallucination detection, most of these approaches rely on computationally expensive sampling strategies and often disregard the distinction between hallucination types. In this work, we introduce a principled evaluation framework that differentiates between extrinsic and intrinsic hallucination categories and evaluates detection performance across a suite of curated benchmarks. In addition, we leverage a recent attention-based uncertainty quantification algorithm and propose novel attention aggregation strategies that improve both interpretability and hallucination detection performance. Our experimental findings reveal that sampling-based methods like Semantic Entropy are effective for detecting extrinsic hallucinations but generally fail on intrinsic ones. In contrast, our method, which aggregates attention over input tokens, is better suited for intrinsic hallucinations. These insights provide new directions for aligning detection strategies with the nature of hallucination and highlight attention as a rich signal for quantifying model uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。