arXiv:2604.00208cs.LG2026-04被引 1

神经网络在叠加态下,传统相似性度量会误判特征相似性。

Similarity of Neural Network Representations in Superposition

论文配图:Similarity of Neural Network Representations in Superposition
图 1 · 摘自论文原文
  • 用闭式推导证明标准度量依赖编码方式而非真实特征
  • 压缩感知下稀疏特征可恢复,但原始激活对齐值偏低
  • 重建潜在特征后对齐度显著提升,适用于真实模型分析

比较内部表示是神经科学和机器学习的核心目标,但标准线性对齐度量(表示相似性分析、中心核对齐、线性回归)常应用于神经活动坐标而非底层特征。当神经系统处于叠加态(通过线性压缩编码超过神经元数量的特征)时,这一做法至关重要。闭式推导表明,这些度量依赖于各系统投影的格拉姆矩阵,而非潜在特征本身:对齐结果混合了系统所表征的内容与编码方式。对于关注两系统共享特征的研究者而言,这构成问题——两个网络可能具有完全相同的特征内容,却表现出比部分特征重叠的网络更不相似。这种看似错位的现象并非信息丢失,因为压缩感知保证稀疏特征仍可从压缩活动中恢复。我们通过训练监督式TopK稀疏自编码器验证此结论,其结构上实现可解压缩感知,发现对恢复的潜在特征进行对齐可恢复对齐度,而原始激活对齐仍偏低。该结果扩展至无监督自编码器及预训练视觉与语言模型的自编码器,其中潜在特征对齐优于原始激活对齐,与真实系统中的叠加现象一致。

原文摘要 · Abstract (English)

Comparing internal representations is a central goal in neuroscience and machine learning, but standard linear alignment metrics (Representational Similarity Analysis, Centered Kernel Alignment, and linear regression) are frequently applied to neural activity coordinates rather than on the underlying features. We show this matters when neural systems operate in superposition, encoding more features than they have neurons via linear compression. Closed-form derivations prove that these metrics depend on the Gram matrices of each system's projection, not on the latent features themselves: alignment thus combines what a system represents with how it is encoded. For those interested in what features two systems share, this is a problem: Two networks can have identical feature content yet appear more dissimilar than networks exhibiting partial feature overlap. This apparent misalignment need not reflect lost information as compressed sensing guarantees sparse features remain recoverable from the compressed activity. We confirm this by training supervised TopK sparse autoencoders that realize solvable compressed sensing by construction, finding alignment on recovered latents restored even when raw-activation alignment remains deflated. We extend the result to unsupervised SAEs trained without ground-truth latents, and to pretrained vision and language model SAEs, where SAE-latent alignment exceeds raw-activation alignment, consistent with superposition in real systems.

神经网络表征叠加态特征对齐压缩感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。