解析神经网络中特征叠加现象的理论与实践,助力模型可解释性研究。
Feature Superposition in Neural Networks: From Theory to Practice

- 从特征统计与解码器选择角度分析叠加表示的几何与学习机制
- 对比多种方法在重建激活时的表现与局限,揭示其对特征识别的影响
- 指出当前方法在因果推理与特征身份确认上的不足,适合模型可解释性研究者
超维度表示指神经网络在维度低于特征数量的情况下仍能编码多个特征,这可能解释多义神经元现象,并为从激活中恢复可解释特征提供思路。理论研究通常基于预设输入特征及其变化规律,分析网络如何在低维隐藏表示中编码这些值;而实证工作则致力于识别训练后网络中实际编码的特征及其计算作用。本文综述了叠加表示的几何、学习与计算特性,阐明特征统计分布与解码器选择对结论的影响。通过比较实用的特征恢复与分析方法,探讨其评估结果所支持的证据。由于精确激活重建不足以证明特征身份或因果使用,我们讨论了现有方法的失败案例与应用情境。最后,评估已有开放问题,并提出关于训练网络中叠加现象的剩余理论与实证挑战。期望本工作能推动对叠加现象的深入理解,并发展更可靠的神经网络解释方法。
原文摘要 · Abstract (English)
Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons and motivates methods for recovering interpretable features from neural activations. Theoretical models typically start with a given set of input features and assumptions about how their values vary across inputs, then study how a network encodes those values in a lower-dimensional hidden representation. Empirical work, by contrast, seeks to identify the features encoded in trained networks and determine their role in computation. In this survey, we review the geometry, learning, and computation of superposed representations, explaining how feature statistics and decoder choice affect the conclusions. To connect these theoretical accounts with evidence from trained networks, we compare practical methods for recovering and analyzing features and examine what their evaluations establish. Since accurate activation reconstruction alone does not establish feature identity or causal use, we discuss the methods' documented failures and applications in light of the evidence available for these different claims. Finally, we assess previously stated open problems and identify remaining theoretical and empirical questions about superposition in trained networks. We hope our work can pave the way for a deeper understanding of superposition and more reliable methods for interpreting neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。