研究稀疏自编码器在特征流形下的扩展规律,发现其可能因流形导致学习效率严重下降。
Understanding sparse autoencoder scaling in the presence of feature manifolds
- 引入容量分配模型分析稀疏自编码器的扩展特性
- 发现特征流形会使实际学习特征数远低于隐层维度
- 提示真实场景中稀疏自编码器可能处于低效状态
稀疏自编码器(SAEs)将神经网络激活值建模为稀疏出现的方向组合(隐变量)。其激活重构能力随隐变量数量呈现缩放规律。本文借鉴神经网络缩放研究中的容量分配模型(Brill, 2024),分析SAE的缩放行为,尤其关注‘特征流形’(多维特征)的影响。结果与先前工作一致,识别出不同缩放阶段。值得注意的是,在某一阶段,特征流形会产生病态效应:SAE学习到的实际特征数远低于其隐变量数量。本文对真实数据中SAE是否处于该病态阶段进行了初步探讨。
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) model the activations of a neural network as linear combinations of sparsely occurring directions of variation (latents). The ability of SAEs to reconstruct activations follows scaling laws w.r.t. the number of latents. In this work, we adapt a capacity-allocation model from the neural scaling literature (Brill, 2024) to understand SAE scaling, and in particular, to understand how "feature manifolds" (multi-dimensional features) influence scaling behavior. Consistent with prior work, the model recovers distinct scaling regimes. Notably, in one regime, feature manifolds have the pathological effect of causing SAEs to learn far fewer features in data than there are latents in the SAE. We provide some preliminary discussion on whether or not SAEs are in this pathological regime in the wild.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。