对比分析了状态空间与Transformer模型的上下文表示流动机制。
A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
- 首次统一分析两种模型在词元和层间的表征传播规律。
- 发现Transformer快速同质化,而状态空间模型早期保持独特性。
- 揭示了两者同质化根源不同,指导未来长序列模型设计。
状态空间模型(SSMs)作为处理长序列的高效替代方案,具有线性计算复杂度,但其上下文信息在各层间的传递机制尚不明确。本文首次对SSMs与基于Transformer的模型(TBMs)进行了统一的、逐词元和逐层的表征传播分析。通过中心核对齐、方差度量及探测方法,刻画了表征在层内与层间的演化过程。研究发现关键差异:TBMs迅速同质化词元表征,多样性仅在深层重新出现;而SSMs早期保持词元独特性,但深层趋于同质化。理论分析与参数随机化进一步表明,TBMs的过度平滑源于架构设计,而SSMs的同质化主要来自训练动态。这些发现揭示了两类架构的归纳偏置,为未来长上下文推理模型与训练策略提供了指导。
原文摘要 · Abstract (English)
State Space Models (SSMs) have recently emerged as efficient alternatives to Transformer-Based Models (TBMs) for long-sequence processing with linear scaling, yet how contextual information flows across layers in these architectures remains understudied. We present the first unified, token- and layer-wise analysis of representation propagation in SSMs and TBMs. Using centered kernel alignment, variance-based metrics, and probing, we characterize how representations evolve within and across layers. We find a key divergence: TBMs rapidly homogenize token representations, with diversity reemerging only in later layers, while SSMs preserve token uniqueness early but converge to homogenization deeper. Theoretical analysis and parameter randomization further reveal that oversmoothing in TBMs stems from architectural design, whereas in SSMs, it arises mainly from training dynamics. These insights clarify the inductive biases of both architectures and inform future model and training designs for long-context reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。