用信息论量化神经网络中特征重叠的压缩程度,揭示其与抗攻击性的关系。
Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability
- 通过稀疏自编码器激活熵计算有效特征数,衡量超叠加带来的压缩效果。
- 发现网络在复杂任务或容量不足时会减少有效特征,而在简单任务中可扩展特征。
- 揭示对抗训练虽提升鲁棒性却可能增加超叠加,挑战传统认知。
神经网络通过超叠加实现卓越性能:将多个特征编码为激活空间中的重叠方向,而非为每个特征分配独立神经元。这虽阻碍可解释性,但缺乏测量超叠加的理论方法。本文提出信息论框架,通过计算稀疏自编码器激活的香农熵,量化表示的有效自由度,即无干扰编码所需的最少神经元数。等价于衡量网络通过超叠加模拟的‘虚拟神经元’数量。当有效特征数超过实际神经元数时,必须以干扰为代价换取压缩。该指标在模型中强相关于真实值,能检测算法任务中最小超叠加,并发现丢弃率下的系统性降低。分层模式与Pythia-70M的内在维度研究一致。该指标还捕捉发展动态,在‘领悟期’检测到特征集中突变。令人意外的是,对抗训练可同时增加有效特征并提升鲁棒性,反驳了‘超叠加导致脆弱性’的假设。实际效果取决于任务复杂度与网络容量:简单任务且容量充足时允许特征扩展(丰裕态),而复杂任务或容量受限则迫使特征缩减(稀缺态)。本工作将超叠加定义为有损压缩,实现了在计算约束下对神经网络信息组织方式的原理性测量,并连接了超叠加与对抗鲁棒性。
原文摘要 · Abstract (English)
Neural networks achieve remarkable performance through superposition: encoding multiple features as overlapping directions in activation space rather than dedicating individual neurons to each feature. This challenges interpretability, yet we lack principled methods to measure superposition. We present an information-theoretic framework measuring a neural representation's effective degrees of freedom. We apply Shannon entropy to sparse autoencoder activations to compute the number of effective features as the minimum neurons needed for interference-free encoding. Equivalently, this measures how many "virtual neurons" the network simulates through superposition. When networks encode more effective features than actual neurons, they must accept interference as the price of compression. Our metric strongly correlates with ground truth in toy models, detects minimal superposition in algorithmic tasks, and reveals systematic reduction under dropout. Layer-wise patterns mirror intrinsic dimensionality studies on Pythia-70M. The metric also captures developmental dynamics, detecting sharp feature consolidation during grokking. Surprisingly, adversarial training can increase effective features while improving robustness, contradicting the hypothesis that superposition causes vulnerability. Instead, the effect depends on task complexity and network capacity: simple tasks with ample capacity allow feature expansion (abundance regime), while complex tasks or limited capacity force reduction (scarcity regime). By defining superposition as lossy compression, this work enables principled measurement of how neural networks organize information under computational constraints, connecting superposition to adversarial robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。