arXiv:2609.08692cs.CL2026-09

对比Transformer与SSM的表征几何,发现结构不同但功能高度趋同。

Global Divergence, Local Convergence: Representation Geometry in SSMs and Transformers

论文配图:Global Divergence, Local Convergence: Representation Geometry in SSMs and Transformers
图 1 · 摘自论文原文
  • SSM表征均匀分布,Transformer则集中于单一主方向。
  • 两者有效表征容量相当,概念编码维度相似。
  • 局部语义流形高度对齐,功能收敛显著。

近期状态空间模型(SSMs)如Mamba在语言建模性能上已接近Transformer,尽管架构本质不同。本研究通过多尺度分析比较了Transformer、SSM及混合架构的内部表征。发现SSM的表征在各维度上均匀分布,而Transformer则高度依赖单一主方向;通过混合架构实验,观察到每经过一层注意力机制,表征空间逐渐向单一方向倾斜。进一步研究表征压缩性发现,尽管几何结构差异明显,二者有效表征容量却紧密匹配。使用秩约束探测器验证,两种架构对概念的编码均发生在维数相近的子空间中,且Transformer的主方向并未蕴含更多概念信息。最后,在局部层面(特定主题或邻近词元)分析流形对齐度,发现两者具有高度一致性。结果表明:虽然两类模型在潜空间使用方式上存在差异,但在局部语义流形层面表现出显著的功能收敛。

原文摘要 · Abstract (English)

Recent state-space models (SSMs) such as Mamba achieve language modeling performance comparable to transformers despite relying on fundamentally different architectures. This raises an important question: how do these structural differences influence the geometry and functional nature of their internal representations? We study this question through a multi-scale analysis of representations in transformers, SSMs, and hybrid architecture. First, we find that SSMs distribute their representational information evenly across all dimensions, whereas transformer representations are heavily dominated by a single principal direction. By evaluating hybrid architectures, we observe that the representation space becomes increasingly skewed toward a single dominant direction after each attention layer. Next, we explore how the different geometric spread of representations impacts representational capacity through compressibility. Surprisingly, we find that despite their contrasting geometric structures, both architectures exhibit tightly matched effective capacities. We further investigate whether this skewed geometry affects how concepts are encoded. Using rank-constrained probes, we demonstrate that both architectures encode concepts in subspaces of surprisingly similar dimensionality. Furthermore, we demonstrate that the transformers' dominant principal direction does not inherently encode more conceptual information. Finally, we zoom in and examine the alignment between manifolds, either by analyzing representations of specific topics or by looking at the nearest neighborhoods of tokens, and find that they are highly aligned. Ultimately, our analysis suggests that while transformers and SSMs induce different usage of latent space, they display a striking functional convergence at the level of local semantic manifolds.

表征几何TransformerSSM功能收敛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。