发现大模型幻觉随上下文累积而增长,关键在注意力机制的漂移。
Shadows in the Attention: Contextual Perturbation and Representation Drift in the Dynamics of Hallucination in LLMs
- 通过渐进注入上下文,追踪隐藏状态与注意力图的动态漂移。
- 5-7轮后幻觉率和表征漂移趋于稳定,关联注意力锁定阈值。
- 相关上下文导致自洽幻觉,无关内容引发注意力重定向错误。
幻觉——看似合理却错误的输出——仍是大型语言模型(LLMs)可靠部署的关键障碍。我们首次系统性地将幻觉发生率与增量上下文注入引发的内部状态漂移联系起来。基于TruthfulQA,为每个问题构建两条16轮‘滴定’轨迹:一条添加相关但部分有误的片段,另一条注入故意误导的内容。在六种开源LLM上,使用三视角检测器追踪明显幻觉率,并通过隐藏状态与注意力图的余弦、熵、JS及斯皮尔曼漂移分析隐性动态。结果表明:(1)幻觉频率与表征漂移呈单调增长,在5–7轮后趋于平稳;(2)相关上下文促进深层语义融合,生成高置信度‘自洽’幻觉;无关上下文则引发由注意力重路由锚定的主题漂移错误;(3)当JS-漂移(≈0.69)与斯皮尔曼-漂移(≈0)收敛时,标志‘注意力锁定’阈值,此后幻觉固化且难以纠正。相关性分析揭示了同化能力与注意力扩散之间的权衡关系,明确了规模依赖的错误模式。这些发现为内在幻觉预测与上下文感知缓解机制提供了实证基础。
原文摘要 · Abstract (English)
Hallucinations -- plausible yet erroneous outputs -- remain a critical barrier to reliable deployment of large language models (LLMs). We present the first systematic study linking hallucination incidence to internal-state drift induced by incremental context injection. Using TruthfulQA, we construct two 16-round "titration" tracks per question: one appends relevant but partially flawed snippets, the other injects deliberately misleading content. Across six open-source LLMs, we track overt hallucination rates with a tri-perspective detector and covert dynamics via cosine, entropy, JS and Spearman drifts of hidden states and attention maps. Results reveal (1) monotonic growth of hallucination frequency and representation drift that plateaus after 5--7 rounds; (2) relevant context drives deeper semantic assimilation, producing high-confidence "self-consistent" hallucinations, whereas irrelevant context induces topic-drift errors anchored by attention re-routing; and (3) convergence of JS-Drift ($\sim0.69$) and Spearman-Drift ($\sim0$) marks an "attention-locking" threshold beyond which hallucinations solidify and become resistant to correction. Correlation analyses expose a seesaw between assimilation capacity and attention diffusion, clarifying size-dependent error modes. These findings supply empirical foundations for intrinsic hallucination prediction and context-aware mitigation mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。