arXiv:2606.18538cs.LGstat.ML2026-06

解析稀疏输入下自编码器的失真机制,揭示超叠加现象的数学原理。

Effects of sparsity and superposition on loss in simple autoencoders

  • 基于幂激活函数建立损失上下界,推导超叠加的数学基础。
  • 在极稀疏条件下,理论损失界与实际性能高度吻合。
  • 为神经元多语义性提供可解释框架,适合研究可解释性者阅读。

神经网络机制可解释性的主要挑战之一是多义性现象,即每个神经元通常负责多个不同任务,阻碍对其功能的清晰解读。Elhage 等(2022)提出,这是由于超叠加现象所致:神经网络在低维空间中以非正交方向表示不同特征,这一策略利用输入向量的特征稀疏性,可在不牺牲保真度的前提下实现更高压缩率。该工作通过一个具有稀疏输入的简单自编码器实证验证了上述假设。本文的贡献在于分析超叠加出现与最优性的数学基础,并严格验证部分发现。具体而言,针对幂激活函数,在极稀疏条件下给出了 L2 重建损失的紧致上下界。文章末尾还列出若干有趣开放问题。

原文摘要 · Abstract (English)

One of the major difficulties in the mechanistic interpretability of neural networks is the occurrence of polysemanticity, which suggests that each neuron is typically responsible for multiple different tasks, impeding a clean interpretation of their function. The seminal paper of Elhage et al. (2022) argues that this occurs due to superposition, a phenomenon where the neural network represents distinct features as non-orthogonal directions in a lower-dimensional space, a strategy that allows much greater compression of the data without sacrificing fidelity due to the feature sparsity of input vectors. Elhage et al. (2022) empirically validates these hypotheses in a rather natural and simple autoencoder with sparse inputs. The contribution of the present work is to analyze the mathematical basis for the occurrence and optimality of superposition, while rigorously corroborating some of their findings. In particular, we provide upper and lower bounds for the L2 reconstruction loss, tight in the very sparse regime, for power activation functions. A short list of interesting open problems are also included at the end.

自编码器可解释性超叠加稀疏性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。