arXiv:2601.07422cs.CLcs.AI2026-01ACL被引 4

揭示大模型幻觉背后的两种内在线索路径,助力提升生成可信度。

Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations

  • 发现幻觉信号来自问答流与答案自证两条独立信息路径。
  • 内部表征能区分两类机制,且与知识边界紧密相关。
  • 基于发现提出新检测方法,可有效提升幻觉识别性能。

尽管大语言模型(LLMs)表现出色,但频繁产生幻觉。以往研究显示其内部状态蕴含丰富的可信度信号,但这些信号的来源与机制尚不明确。本文证明,可信度线索源于两条不同信息路径:(1) 依赖问答信息流动的‘问题锚定’路径;(2) 从生成答案自身提取自洽证据的‘答案锚定’路径。通过注意力掩码与标记修补实验,我们验证并解耦了这两条路径。进一步实验揭示:(1) 两种机制与模型知识边界密切相关;(2) 内部表示能意识到二者差异。基于此,我们提出了两项应用以提升幻觉检测性能。整体而言,本工作为理解大模型如何内嵌可信度提供了新视角,推动更可靠、自知的生成系统发展。

原文摘要 · Abstract (English)

Despite their impressive capabilities, large language models (LLMs) frequently generate hallucinations. Previous work shows that their internal states encode rich signals of truthfulness, yet the origins and mechanisms of these signals remain unclear. In this paper, we demonstrate that truthfulness cues arise from two distinct information pathways: (1) a Question-Anchored pathway that depends on question-answer information flow, and (2) an Answer-Anchored pathway that derives self-contained evidence from the generated answer itself. First, we validate and disentangle these pathways through attention knockout and token patching. Afterwards, we uncover notable and intriguing properties of these two mechanisms. Further experiments reveal that (1) the two mechanisms are closely associated with LLM knowledge boundaries; and (2) internal representations are aware of their distinctions. Finally, building on these insightful findings, two applications are proposed to enhance hallucination detection performance. Overall, our work provides new insight into how LLMs internally encode truthfulness, offering directions for more reliable and self-aware generative systems.

大模型幻觉检测可信度内部机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。