通过权重符号结构分析,揭示神经网络计算布尔函数的内在机制。
Towards Combinatorial Interpretability of Neural Computation
- 基于权重符号构建特征通道编码,解析神经元的组合计算逻辑。
- 首次实现无需激活或重训练,静态分析即解释小网络的完整计算过程。
- 为理解人工与生物神经回路提供可量化的组合解释新范式。
我们提出组合可解释性方法,通过分析网络权重和偏置的符号分类中的组合结构来理解神经计算。通过特征通道编码理论,该方法解释了神经网络如何计算布尔表达式,并可能揭示其他类型的神经网络计算。特征通过特征通道实现:共享输入间跨神经元的独特编码。由于不同通道共享神经元,神经元具有多义性,通道间相互干扰,使计算看似不可理解。我们通过分析特征通道编码,成功解码多个梯度下降训练的小型神经网络的计算过程,且仅依赖权重矩阵的静态组合分析,无需考察激活值或训练新自编码网络。特征通道编码重构了超位置假设,将关注点从高维空间中神经元激活方向转向编码的组合结构。它还首次精确量化并解释了网络参数规模与计算能力(即低误差计算特征集)之间的关系,这一关系隐含于众多现代缩放定律中。尽管当前研究局限于布尔函数,但我们认为其提供了丰富、可控且富有信息的研究空间,所提出的组合解释路径有望为理解人工与生物神经电路奠定基础。
原文摘要 · Abstract (English)
We introduce combinatorial interpretability, a methodology for understanding neural computation by analyzing the combinatorial structures in the sign-based categorization of a network's weights and biases. We demonstrate its power through feature channel coding, a theory that explains how neural networks compute Boolean expressions and potentially underlies other categories of neural network computation. According to this theory, features are computed via feature channels: unique cross-neuron encodings shared among the inputs the feature operates on. Because different feature channels share neurons, the neurons are polysemantic and the channels interfere with one another, making the computation appear inscrutable. We show how to decipher these computations by analyzing a network's feature channel coding, offering complete mechanistic interpretations of several small neural networks that were trained with gradient descent. Crucially, this is achieved via static combinatorial analysis of the weight matrices, without examining activations or training new autoencoding networks. Feature channel coding reframes the superposition hypothesis, shifting the focus from neuron activation directionality in high-dimensional space to the combinatorial structure of codes. It also allows us for the first time to exactly quantify and explain the relationship between a network's parameter size and its computational capacity (i.e. the set of features it can compute with low error), a relationship that is implicitly at the core of many modern scaling laws. Though our initial studies of feature channel coding are restricted to Boolean functions, we believe they provide a rich, controlled, and informative research space, and that the path we propose for combinatorial interpretation of neural computation can provide a basis for understanding both artificial and biological neural circuits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。