arXiv:2603.01326cs.CLcs.LG2026-03ACL被引 9

通过分析模型推理轨迹,揭示大模型真实推理路径。

Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning

  • 将模型推理视为层间表征的动态演化轨迹
  • 发现几何不变性可区分有效与虚假推理
  • 适用于密集和MoE架构,提升解释可靠性

现有大语言模型可解释性方法通常将隐藏状态视为激活空间中的静态点,假设正确与错误推理可通过单层表示区分。然而这些激活富含多义特征,导致线性探测器学习表面词汇模式而非底层推理结构。我们提出真理即轨迹(TaT),将Transformer推理建模为逐层迭代优化的展开轨迹,将分析从静态激活转向层间几何位移。通过分析跨层表示的位移,TaT揭示了区分有效推理与虚假行为的几何不变性。我们在涵盖常识推理、问答和毒性检测的基准上评估了TaT,无需访问原始激活,仅使用层间变化,便有效缓解对静态词汇混淆的依赖,优于传统探测方法,确立轨迹分析作为大模型可解释性的互补视角。

原文摘要 · Abstract (English)

Existing explainability methods for Large Language Models (LLMs) typically treat hidden states as static points in activation space, assuming that correct and incorrect inferences can be separated using representations from an individual layer. However, these activations are saturated with polysemantic features, leading to linear probes learning surface-level lexical patterns rather than underlying reasoning structures. We introduce Truth as a Trajectory (TaT), which models the transformer inference as an unfolded trajectory of iterative refinements, shifting analysis from static activations to layer-wise geometric displacement. By analyzing displacement of representations across layers, TaT uncovers geometric invariants that distinguish valid reasoning from spurious behavior. We evaluate TaT across dense and Mixture-of-Experts (MoE) architectures on benchmarks spanning commonsense reasoning, question answering, and toxicity detection. Without access to the activations themselves and using only changes in activations across layers, we show that TaT effectively mitigates reliance on static lexical confounds, outperforming conventional probing, and establishes trajectory analysis as a complementary perspective on LLM explainability.

大模型解释推理轨迹几何不变性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。