arXiv:2603.06557cs.LGq-bio.NC2026-03被引 2

通过分解神经元贡献,揭示网络决策背后的因果机制。

Causal Interpretation of Neural Network Computations with Contribution Decomposition

  • 用稀疏自编码器将隐藏层贡献拆解为稀疏模式
  • 发现贡献随层数增加更稀疏、正负效应逐步解耦
  • 可实现对中间层的精准操控与人类可读的图像成分可视化

理解神经网络如何将输入转化为输出,对解释和操控其行为至关重要。现有方法多通过分析与可解释概念相关的隐藏层激活模式来实现。本文提出直接考察隐藏神经元如何驱动网络输出的方法——CODEC(贡献分解),利用稀疏自编码器将网络行为分解为稀疏的神经元贡献模式,揭示仅靠激活分析无法获得的因果过程。在基准图像分类模型上应用CODEC发现,贡献随层加深逐渐变得更稀疏且维度更高,并意外地表现出正负效应逐步解耦。进一步表明,将贡献分解为稀疏模式可增强对中间层的控制与解释能力,支持对输出的因果干预以及对驱动输出的图像组件进行人类可读的可视化。最后,在脊椎动物视网膜神经活动的先进模型中,CODEC揭示了模型中间神经元的组合行为并定位动态感受野的来源。总体而言,CODEC提供了一个丰富且可解释的框架,用于理解非线性计算在分层结构中的演化,确立贡献模式作为解析人工神经网络机制的有力分析单元。

原文摘要 · Abstract (English)

Understanding how neural networks transform inputs into outputs is crucial for interpreting and manipulating their behavior. Most existing approaches analyze internal representations by identifying hidden-layer activation patterns correlated with human-interpretable concepts. Here we take a direct approach to examine how hidden neurons act to drive network outputs. We introduce CODEC (Contribution Decomposition), a method that uses sparse autoencoders to decompose network behavior into sparse motifs of hidden-neuron contributions, revealing causal processes that cannot be determined by analyzing activations alone. Applying CODEC to benchmark image-classification networks, we find that contributions grow in sparsity and dimensionality across layers and, unexpectedly, that they progressively decorrelate positive and negative effects on network outputs. We further show that decomposing contributions into sparse modes enables greater control and interpretation of intermediate layers, supporting both causal manipulations of network output and human-interpretable visualizations of distinct image components that combine to drive that output. Finally, by analyzing state-of-the-art models of neural activity in the vertebrate retina, we demonstrate that CODEC uncovers combinatorial actions of model interneurons and identifies the sources of dynamic receptive fields. Overall, CODEC provides a rich and interpretable framework for understanding how nonlinear computations evolve across hierarchical layers, establishing contribution modes as an informative unit of analysis for mechanistic insights into artificial neural networks.

神经网络解释因果分析贡献分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。