通过激活谱分析,揭示神经网络中分布式表示的协同神经元组合。
Making Sense Of Distributed Representations With Activation Spectroscopy
- 将层间激活模式建模为伪布尔函数,用傅里叶系数量化神经元子集贡献。
- 提出组合优化算法,找出高值且非冗余的傅里叶系数,识别关键神经元组。
- 适用于理解图像分类与情感分析模型中隐藏的分布式特征编码机制。
在神经网络可解释性研究中,越来越多证据表明相关特征以分布式方式编码于多个神经元中。在不了解网络编码策略的情况下,解析这些分布式表示是一个组合难题,未必可解。本文提出一种可行路径——激活谱分析(ActSpec),通过分析网络层激活模式定义的伪布尔傅里叶谱,检测并追踪分布式表示中神经元的联合影响。将给定层与输出逻辑单元之间的子网络视为一类特殊伪布尔函数,其傅里叶系数可量化各神经元子集的贡献。我们设计了一种组合优化方法,用于搜索同时具有高值且非冗余的傅里叶系数,该方法可看作引入特定约束的Goldreich-Levin算法扩展。所得系数确定一组神经元子集,用于评估表示的分布程度。我们在多个合成场景下验证方法,并与现有可解释性基准对比。最后,在一个MNIST分类器和一个基于Transformer的情感分析网络上进行实验评估。
原文摘要 · Abstract (English)
In the study of neural network interpretability, there is growing evidence to suggest that relevant features are encoded across many neurons in a distributed fashion. Making sense of these distributed representations without knowledge of the network's encoding strategy is a combinatorial task that is not guaranteed to be tractable. This work explores one feasible path to both detecting and tracing the joint influence of neurons in a distributed representation. We term this approach Activation Spectroscopy (ActSpec), owing to its analysis of the pseudo-Boolean Fourier spectrum defined over the activation patterns of a network layer. The sub-network defined between a given layer and an output logit is cast as a special class of pseudo-Boolean function. The contributions of each subset of neurons in the specified layer can be quantified through the function's Fourier coefficients. We propose a combinatorial optimization procedure to search for Fourier coefficients that are simultaneously high-valued, and non-redundant. This procedure can be viewed as an extension of the Goldreich-Levin algorithm which incorporates additional problem-specific constraints. The resulting coefficients specify a collection of subsets, which are used to test the degree to which a representation is distributed. We verify our approach in a number of synthetic settings and compare against existing interpretability benchmarks. We conclude with a number of experimental evaluations on an MNIST classifier, and a transformer-based network for sentiment analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。