arXiv:2510.06107cs.CLcs.AI2025-10被引 3

通过语义追踪揭示大模型幻觉背后的深层机制

Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models

  • 基于残差流构建逐层语义图,定位生成位置的关联概念
  • 发现幻觉源于训练共现导致的语义漂移,与上下文不一致
  • 在赛车思维数据集上表现优于基线方法,适合模型可解释性研究

大语言模型在上下文信息不足或模糊时会产生流畅但无依据的输出,即幻觉。我们提出分布语义追踪(DST),一种原生模型方法:通过解码残差流状态,经由反嵌入选择紧凑的前K个核心概念,并利用轻量级因果追踪估计概念间的定向支持关系,构建答案位置的逐层语义映射。基于这些追踪结果,我们检验了一个表示层面的假设:幻觉源于深度方向上的相关性驱动的表征漂移,即残差流被拉向一个局部连贯但与上下文矛盾的概念邻域,该现象由训练中的共现强化所致。在Racing Thoughts数据集上,DST在模型裁判协议下生成的解释比归因、探针和干预基线更准确,且得到的上下文对齐得分(CAS)能强预测失败,支持该漂移假说。

原文摘要 · Abstract (English)

Hallucinations in large language models (LLMs) produce fluent continuations that are not supported by the prompt, especially under minimal contextual cues and ambiguity. We introduce Distributional Semantics Tracing (DST), a model-native method that builds layer-wise semantic maps at the answer position by decoding residual-stream states through the unembedding, selecting a compact top-$K$ concept set, and estimating directed concept-to-concept support via lightweight causal tracing. Using these traces, we test a representation-level hypothesis: hallucinations arise from correlation-driven representational drift across depth, where the residual stream is pulled toward a locally coherent but context-inconsistent concept neighborhood reinforced by training co-occurrences. On Racing Thoughts dataset, DST yields more faithful explanations than attribution, probing, and intervention baselines under an LLM-judge protocol, and the resulting Contextual Alignment Score (CAS) strongly predicts failures, supporting this drift hypothesis.

大模型幻觉语义追踪可解释性因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。