解开神经网络的特征叠加,能更真实地揭示模型间的信息对齐程度。
Superposition disentanglement of neural representations reveals hidden alignment
- 通过稀疏自编码器解耦特征叠加,使神经元表示更清晰。
- 解耦后模型间的线性映射对齐分数普遍提升,最高增益达23%。
- 适用于研究深度神经网络与大脑视觉表征对齐的研究者。
超叠加假说认为单个神经元可参与多个特征的表征,使网络能表达超过其神经元数量的特征。在神经科学与人工智能中,表征对齐度量用于评估不同深度神经网络(DNNs)或大脑是否表征相似信息。本文探讨关键问题:超叠加是否对对齐度量产生不利影响?我们假设,若模型以不同超叠加方式表示相同特征(即神经元采用不同特征的线性组合),则会干扰预测映射度量(半匹配、软匹配、线性回归),导致对齐度低于预期。我们建立了排列度量依赖于超叠加结构的理论,并通过训练稀疏自编码器(SAEs)在模拟模型中解耦超叠加,发现将基底神经元替换为稀疏超完备潜在码后,对齐得分通常显著提升。该现象在视觉领域的DNN-DNN与DNN-脑线性回归对齐中均被观察到。结果表明,超叠加解耦是揭示模型间真实表征对齐的必要条件。
原文摘要 · Abstract (English)
The superposition hypothesis states that single neurons may participate in representing multiple features in order for the neural network to represent more features than it has neurons. In neuroscience and AI, representational alignment metrics measure the extent to which different deep neural networks (DNNs) or brains represent similar information. In this work, we explore a critical question: does superposition interact with alignment metrics in any undesirable way? We hypothesize that models which represent the same features in different superposition arrangements, i.e., their neurons have different linear combinations of the features, will interfere with predictive mapping metrics (semi-matching, soft-matching, linear regression), producing lower alignment than expected. We develop a theory for how permutation metrics are dependent on superposition arrangements. This is tested by training sparse autoencoders (SAEs) to disentangle superposition in toy models, where alignment scores are shown to typically increase when a model's base neurons are replaced with its sparse overcomplete latent codes. We find similar increases for DNN-DNN and DNN-brain linear regression alignment in the visual domain. Our results suggest that superposition disentanglement is necessary for mapping metrics to uncover the true representational alignment between neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。