arXiv:2507.06722cs.CLcs.LG2025-07中稿 · ICML被引 2

发现大模型推理时不确定性不影响概率变化轨迹

On the Effect of Uncertainty on Layer-wise Inference Dynamics

  • 用改进的Logit Lens分析各层输出概率变化
  • 确定与不确定预测在相似层出现信心突增
  • 提示简单方法难检测不确定性,适合模型解释研究者

理解大语言模型如何内部表征和处理预测是检测不确定性与防止幻觉的核心。尽管已有研究显示模型在隐藏状态中编码不确定性,但其对隐藏状态处理方式的影响仍不明确。本文通过使用Tuned Lens(Logit Lens的一种变体),分析了11个数据集和5种模型中最终预测标记的逐层概率轨迹。以错误预测作为高认知不确定性样本,结果表明:确定性与不确定性预测的概率轨迹高度一致,均在相近层出现信心突然上升。我们补充指出,更优模型可能以不同方式处理不确定性。该发现挑战了在推理阶段使用简单方法检测不确定性的可行性。更广泛地,本工作展示了可解释性方法在探究不确定性如何影响推理过程中的应用价值。

原文摘要 · Abstract (English)

Understanding how large language models (LLMs) internally represent and process their predictions is central to detecting uncertainty and preventing hallucinations. While several studies have shown that models encode uncertainty in their hidden states, it is underexplored how this affects the way they process such hidden states. In this work, we demonstrate that the dynamics of output token probabilities across layers for certain and uncertain outputs are largely aligned, revealing that uncertainty does not seem to affect inference dynamics. Specifically, we use the Tuned Lens, a variant of the Logit Lens, to analyze the layer-wise probability trajectories of final prediction tokens across 11 datasets and 5 models. Using incorrect predictions as those with higher epistemic uncertainty, our results show aligned trajectories for certain and uncertain predictions that both observe abrupt increases in confidence at similar layers. We balance this finding by showing evidence that more competent models may learn to process uncertainty differently. Our findings challenge the feasibility of leveraging simplistic methods for detecting uncertainty at inference. More broadly, our work demonstrates how interpretability methods may be used to investigate the way uncertainty affects inference.

大模型不确定性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。