发现预测方向是模型残差流中的关键锚点,决定信息组织方式。
Geometric and Behavioral Stratification in Transformer Residual Streams

- 以预测方向为锚,残差流呈现几何与行为分层结构。
- 近预测区域高度有序,远预测区域平坦且反区分提示组。
- 该结构在18个模型中一致存在,对线性读出有决定性作用。
训练好的Transformer模型会发展出特殊坐标系:其坐标轴的统计特性不同于残差流其余部分。我们研究了模型当前预测的标记的预测方向(即解嵌入方向),发现它充当了一个由内容定义的特权锚点。以该锚点为基准,残差流的变异呈现出几何与行为上的分层——越靠近预测方向,结构越明显。这一分层现象在全部18个测试模型中成立(包括密集和专家混合模型,参数量7B至120B,基础与指令微调版本)。一个狭窄且尺度不变的预测接口集中了读出相关结构,而远离预测的部分随模型规模扩大而扩展。由于预测方向几乎正交于主方差轴,基于方差的分析仅部分揭示该结构,且在提示异质性越高时差距越大。以锚点为参照,可观察到陡峭的几何梯度:靠近预测的区域高度结构化并聚类相关提示,而远端则更平坦、反区分提示组。该接口虽仅为窄切片,却功能决定性。破坏最接近预测的方差方向会导致立即发散和频繁任务框架变化;破坏次一级方向则延迟发散并保持框架稳定。远端部分虽每方向读出对齐弱,但因果和时间上承担重要角色,行为由方向而非幅度驱动。这些结果确立预测方向为独立于先前描述坐标的特权锚点,并为高维计算如何与线性读出共存提供了几何解释。
原文摘要 · Abstract (English)
Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. But what kind of direction does such a basis select? We investigate the prediction direction, the unembedding direction of the token a model currently predicts, and find that it functions as a content-defined privileged anchor. Measured with respect to this anchor, residual-stream variation is geometrically and behaviorally stratified by proximity to the prediction. The stratification holds in all eighteen models tested (dense and mixture-of-experts, 7B-120B, base and instruction-tuned). A narrow, scale-invariant prediction interface concentrates readout-relevant structure, while the vast prediction-distal complement expands with model scale. Because the prediction direction sits nearly orthogonal to the principal variance axes, variance-based analyses recover this organization only partly, and the shortfall grows with prompt heterogeneity. Anchoring reveals a steep geometric gradient: prediction-proximal regions are highly structured and cluster related prompts, while the complement is flatter and anti-discriminates among prompt groups. The interface is a narrow slice but functionally decisive. Disrupting the variance directions closest to the prediction causes immediate divergence and frequent task-frame shifts; disrupting the next level down delays divergence and preserves framing. The complement is weakly readout-aligned per direction yet causally and temporally load-bearing, and behavior is driven by direction rather than magnitude. These results establish the prediction direction as a privileged anchor distinct from previously described coordinate axes, and give a geometric account of how high-dimensional computation coexists with linear readout.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。