揭示视觉Transformer中语义结构随层数演化的几何规律。
Transformer Geometry Observatory TGO-III: Semantic Geometry Observatory

- 构建多观测器框架,量化特征演化与类别几何组织。
- 类中心间距增大,判别性提升,局部流形具结构化复杂性。
- 验证语义扩展假说,适合研究模型表征与训练动态者阅读。
随着视觉Transformer在现代AI中的广泛应用,分析其内在表征行为变得日益重要。尽管现有研究多聚焦于标记几何与训练动态,但表征协方差结构及类别级几何组织的演化仍相对未被充分探索。本文通过TGO-III:语义几何观测仪,研究ViT-Small/16各层中语义几何与类别可分性随训练的演变。该框架整合线性探测准确率、费雪比、类别中心距、局部内在维数和局部PCA秩等多重互补观测器,量化判别性表征的渐进演化。结果表明,类别表示逐渐更易线性分离,费雪判别力增强,类别中心距离扩大,局部表征流形呈现具有类别依赖性的结构化复杂性。这些发现为语义扩展假说提供了实证支持,即此前观测到的流形扩张伴随表征向更具判别性的语义结构逐步组织。TGO-III通过建立流形几何、协方差演化与语义组织间的直接关联,扩展了Transformer几何观测框架。
原文摘要 · Abstract (English)
With the widespread adoption of Vision Transformers in modern AI, the need to analyze their inherent representational behavior has become increasingly important. While most existing studies emphasize token geometries and training dynamics, the evolution of representational covariance structures and class-level geometric organization remains comparatively underexplored. In this work, we investigate semantic geometry and class separability as representations evolve across the layers of ViT-Small/16 through TGO-III: Semantic Geometry Observatory. It is a framework designed to analyze the emergence of semantic organization, feature evolution, and class-wise representation geometry throughout training. The framework employs multiple complementary observatories, including Linear Probe Accuracy, Fisher Ratio, Class Centroid Distances, Local Intrinsic Dimension, and Local PCA Rank, to quantify the progressive evolution of discriminative representations. Our analysis reveals that class representations become progressively more linearly separable, Fisher discriminability increases, class centroids move farther apart, and local representation manifolds exhibit structured class-dependent geometric complexity. These observations provide empirical evidence supporting the Semantic Expansion Hypothesis, suggesting that the manifold expansion observed in previous observatories is accompanied by the progressive organization of representations into increasingly discriminative semantic structures. Collectively, TGO-III extends the Transformer Geometry Observatory framework by establishing a direct connection between manifold geometry, covariance evolution, and semantic organization during Transformer training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。