arXiv:2605.09887cs.LGcs.AI2026-05

发现稀疏自编码器的性能瓶颈由激活空间几何结构决定,而非资源限制。

The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws

论文配图:The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws
图 1 · 摘自论文原文
  • 通过跨层分析,揭示激活流形的曲率与内在维度影响稀疏编码器的缩放规律。
  • 不同层的编码器宽度指数与流形几何高度相关,且在模型间可迁移。
  • 性能下限由流形弯曲程度决定,是线性逼近无法消除的固有残差。

稀疏自编码器(SAEs)基于线性表示假设,将模型激活重构为可解释字典原子的稀疏线性组合,隐含假设激活空间具有全局线性结构。然而,现有缩放定律无法解释各层重建误差的剧烈变化。本文认为这是几何失配的实证表现:当激活流形存在曲率且内在维度随层变化时,单一稀疏线性字典无法统一拟合,导致缩放规律变为依赖于流形结构的层特定函数。我们首次开展跨层SAE缩放研究,在Gemma 2 2B和9B共68层的844个残差流检查点上进行建模。第一阶段拟合各层缩放律表面;第二阶段将参数与推导出的每层宽度指数对四个层级几何指标回归。结果表明,流形几何能预测两模型的每层宽度指数,且在一个模型上学到的回归系数可准确预测另一模型的指数,显示可迁移的几何规律。在允许识别渐近下界的展示层中,拟合的下界与层级几何顺序一致:更高曲率与更高内在维度对应更高下界,符合任何稀疏线性逼近曲面流形时必然存在的二阶残差。因此,SAEs面临的并非资源上限,而是由其试图重构的流形决定的几何之墙。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) operationalise the linear representation hypothesis: they reconstruct model activations as sparse linear combinations of interpretable dictionary atoms, on the implicit assumption that activation space is well approximated by a globally linear structure. Their reconstruction error varies sharply across layers in ways that existing scaling laws, fitted at single layers, do not explain. We argue that this variation is the empirical trace of a geometric mismatch: where the activation manifold is curved and its intrinsic dimension varies across layers, no sparse linear dictionary can match it uniformly, and the SAE's width-sparsity scaling becomes a layer-dependent function of manifold structure rather than a single universal law. We conduct the first cross-layer SAE scaling study, fitting and regressing on 844 residual-stream Gemma Scope SAE checkpoints across 68 layers of Gemma 2 2B and 9B. Stage 1 fits a per-layer scaling-law surface; Stage 2 regresses the fitted parameters and the derived per-layer width exponents on four layerwise geometric summaries. We find that manifold geometry predicts the per-layer width exponent in both models, and that the same regression coefficients learnt on one model predict the other model's per-layer exponents under cross-model transfer, indicating a transferable geometric law. At the showcase layers where richer width grids permit identification of the asymptotic floor, we find that the fitted floor tracks the layerwise geometric ordering: higher curvature and intrinsic dimension correspond to higher floor, consistent with the irreducible second-order residual that any sparse linear approximation of a curved manifold must leave behind. SAEs thus encounter not a finite-resource ceiling but a geometry-dependent wall, set by the manifold they are trying to reconstruct.

稀疏编码流形学习缩放律模型解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。