arXiv:2606.12138cs.LGcs.AI2026-06

揭示稀疏自编码器中不稳定特征的本质:非噪声,而是可复现的低维结构。

Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders

论文配图:Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders
图 1 · 摘自论文原文
  • 通过种子依赖性分析,区分稳定与不稳定的特征
  • 稳定特征承载主要重构与预测信号,不稳定的则影响微弱
  • 不稳定的特征集中于可复现的低秩子空间,适合研究表征结构

稀疏自编码器(SAEs)广泛用于解析神经网络表征,但其有效性取决于特征在不同训练运行中是否可复现。本文通过特征稳定性分析——估计每个特征在独立训练的SAE中再次出现的概率——获得一种可扩展的逐特征信号,以区分稳定与不稳定特征。大规模实验覆盖不同种子、模型、层、字典大小和SAE变体,发现显著的功能不对称:稳定特征携带大部分重构与预测相关信号,而不稳定特征边际影响弱,主要受低频表面形式触发,在激活统计与自动解释中占主导。几何上,不稳定特征虽个体不可复现,却集中在可复现的低秩子空间中,表明种子依赖性常反映激活空间共享区域内基的模糊性,而非纯粹噪声。一个可控合成模型明确展示了这一机制:低秩真实特征可在子空间层面被恢复,但在单个种子间无法作为独立潜变量识别。最后,通过合并跨种子唯一特征,构建出更稳定的SAE,同时保持解释方差。结果表明,不稳定特征并非失败或噪声,而是具有弱个体功能影响,却反映可复现的低维结构。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on whether the learned features are reproducible across training runs. We study this question through \emph{feature stability}: for each SAE feature, we estimate the probability that a similar feature reappears in an independently trained SAE. This yields a scalable per-feature signal that separates stable from unstable features. In a large-scale study across seeds, models, layers, dictionary sizes, and SAE variants, we find a pronounced functional asymmetry: stable features carry most of the reconstruction- and prediction-relevant signal, while unstable features have weak marginal impact and are dominated by low-frequency surface-form triggers in both activation statistics and automatic explanations. Geometrically, unstable features are individually non-reproducible but concentrate in reproducible lower-rank subspaces, suggesting that seed dependence often reflects basis ambiguity within a shared region of activation space rather than pure noise. A controlled synthetic model makes this mechanism explicit, showing that low-rank ground-truth features can be recovered at the subspace level while remaining non-identifiable as individual SAE latents across seeds. Finally, by pooling unique cross-seed features, we construct more stable SAEs while preserving explained variance in this setting. Together, these results show that unstable features are not merely failed or noisy latents: they have weak individual functional impact, but reflect reproducible low-dimensional structure that standard SAEs resolve differently across seeds.

稀疏自编码器特征可复现性低秩结构种子依赖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。