arXiv:2606.19249cs.CVcs.LG2026-06被引 2

揭示视觉Transformer的谱几何演化规律,发现训练中维度利用率持续上升

Transformer Geometry Observatory TGO-I: Spectral Geometry Observatory

论文配图:Transformer Geometry Observatory TGO-I: Spectral Geometry Observatory
图 1 · 摘自论文原文
  • 构建TGO框架,分析ViT在ImageNet-100上训练全过程的谱几何特性
  • 训练中有效维度上升、各向异性下降,特征分布趋向更平坦
  • 最终分类令牌表现最高维度与最低各向异性,反直觉但关键

尽管视觉变换器(ViTs)在众多计算机视觉任务中广泛应用并取得成功,其维度与表征几何的基本理解仍相对不足。为填补这一空白,我们提出变换器几何观测仪(TGO),一个系统性的实验与分析流程框架,用于研究视觉变换器的表征几何与动态特性。TGO-I是该框架的首个版本,聚焦于ViT表征的谱几何。基于在ImageNet-100上训练的ViT-Small/16模型,我们分析了有效秩、稳定秩、参与度比、谱熵、谱平坦度、谱各向异性、协方差结构、特征值谱和奇异值谱在整个训练过程中的变化。结果表明,维度利用率持续上升,各向异性逐渐降低,谱熵和参与度比增加,特征值谱趋于更平坦。与常见直觉相反,训练并未将信息集中在少数主导方向,而是促使方差在表征维度间逐步重新分布。这一现象在最后的CLS token表征中尤为显著,其表现出网络中最高的有效维度与最低的各向异性。

原文摘要 · Abstract (English)

Despite the widespread adoption of Vision Transformers (ViTs) and their success across numerous computer vision applications, the fundamental understanding of their dimensional and representational geometry remains relatively underexplored. To address this gap, we introduce Transformer Geometry Observatory (TGO), a systematic framework of experiments and analysis pipelines designed to investigate the representational geometry and dynamics of Vision Transformers. TGO-I, the first installment of the framework, focuses on the spectral geometry of ViT representations. Using a ViT-Small/16 model trained on ImageNet-100, we analyze Effective Rank, Stable Rank, Participation Ratio, Spectral Entropy, Spectral Flatness, Spectral Anisotropy, covariance structure, eigenspectra, and singular value spectra throughout training. Our results reveal a consistent increase in dimensional utilization, accompanied by decreasing anisotropy, increasing spectral entropy, increasing participation ratio, and progressively flatter eigenspectra. Contrary to the common intuition that training should concentrate information into a small number of dominant directions, we observe a progressive redistribution of variance across representational dimensions. This phenomenon is particularly pronounced in the final CLS token representation, which exhibits the highest effective dimensionality and lowest anisotropy within the network.

视觉Transformer谱几何表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。