arXiv:2606.09287cs.LG2026-06

通过轨迹几何分析Transformer层间表示演化,揭示其内在计算机制。

Trajectory Geometry of Transformer Representations Across Layers

论文配图:Trajectory Geometry of Transformer Representations Across Layers
图 1 · 摘自论文原文
  • 将Transformer前向传播视为高维空间中的轨迹,用五项几何指标直接分析
  • 推理任务轨迹曲率更高(0.71–0.83弧度),体现计算复杂性差异
  • 模糊词元在末层轨迹分叉,分离度达5.6倍,适合研究模型不确定性

理解Transformer表示在层间如何演变,而不仅仅是它们编码了什么,仍是机制可解释性中的开放问题。我们将Transformer前向传播重新建模为高维表示流形上的离散群体轨迹,借鉴计算神经科学的几何工具。不依赖预设特征探测,而是直接在环境空间中计算五项指标:轨迹长度、曲率、语义收敛指数(CI)、逐层余弦相似度和表征稳定性。在三个模型族(GPT-2、TinyLlama、Qwen2.5)和五个受控提示族中,我们报告四项发现:第一,语义相关提示在中后期层显著收敛(峰值CI 0.41–0.58,p<0.001,Mann-Whitney U),符合吸引子动态;第二,推理任务轨迹曲率高于词汇变化(0.71–0.83弧度 vs. 0.27–0.31弧度),表明曲率编码计算复杂性;第三,模糊词元表现出轨迹分叉,末层表征分离度最高达5.6倍,而无歧义对照组中无此现象;第四,逐层余弦相似度揭示出统一的三阶段结构:编码、展开与输出准备,跨三类架构一致。所有效应在层随机打乱和随机嵌入控制下均消失。我们开源了完全免探针、模型无关的分析管道,并主张轨迹几何构成一种原则性强、无需探测的机制可解释性视角。

原文摘要 · Abstract (English)

Understanding how transformer representations evolve across layers, not merely what they encode, remains an open problem in mechanistic interpretability. We recast the transformer forward pass as a discrete population trajectory through a high-dimensional representation manifold, drawing on geometric tools from computational neuroscience. Rather than probing for pre-specified features, we characterize trajectory geometry using five metrics computed directly in the ambient space: trajectory length, curvature, a semantic convergence index, layerwise cosine similarity, and representational stability. Across three model families (GPT-2, TinyLlama, Qwen2.5) and five controlled prompt families, we report four findings. First, semantically related prompts converge significantly in middle-to-late layers (peak CI 0.41--0.58, p<0.001, Mann-Whitney U), consistent with attractor-like dynamics. Second, reasoning tasks produce trajectories of greater curvature than lexical variations (0.71--0.83 rad vs. 0.27--0.31 rad), suggesting curvature encodes computational complexity. Third, ambiguous tokens exhibit trajectory bifurcation with up to 5.6x representational separation by the final layer, absent in unambiguous controls. Fourth, layerwise cosine similarity reveals a universal three-phase structure: encoding, elaboration, and output preparation, consistent across all three architectures. All four effects vanish under shuffled-layer and random-embedding controls. We release a fully open-source, model-agnostic pipeline and argue that trajectory geometry constitutes a principled, probe-free lens for mechanistic interpretability.

Transformer轨迹几何可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。