用控制理论解析神经网络内部机制,揭示关键神经元与路径重要性。
From Black-Box to White-Box: Control-Theoretic Neural Network Interpretability
- 将神经网络视为非线性动态系统,通过局部线性化构建状态空间模型。
- 计算可控性、可观测性矩阵和汉克尔奇异值,量化神经元与路径重要性。
- 可识别可剪枝方向,适用于提升模型可解释性的研究者。
深度神经网络虽表现优异,但内部机制难以理解。本文提出一种控制理论框架,将训练好的神经网络视为非线性状态空间系统,通过局部线性化、可控性与可观测性格拉米安矩阵以及汉克尔奇异值分析其内部计算过程。针对特定输入,在对应隐藏激活模式附近对网络进行线性化,构建以隐藏神经元激活为状态的状态空间模型。输入-状态与状态-输出雅可比矩阵定义局部可控性与可观测性格拉米安,进而计算汉克尔奇异值及其关联模态。这些量提供了一种原则化的神经元与路径重要性度量:可控性反映输入扰动激发神经元的难易程度,可观测性反映神经元对输出的影响强度,汉克尔奇异值按输入输出能量排序内部模态。我们在简单前馈网络(如1-2-2-1 SwiGLU和2-3-3-2 GELU)上验证该框架。通过对比不同工作点,发现激活饱和会降低可控性,缩小主导汉克尔奇异值,并使主导内部模态转移至不同的神经元子集。该方法将神经网络转化为一系列局部白盒动态模型,指明了适合剪枝或约束的内部方向,以提升可解释性。
原文摘要 · Abstract (English)
Deep neural networks achieve state of the art performance but remain difficult to interpret mechanistically. In this work, we propose a control theoretic framework that treats a trained neural network as a nonlinear state space system and uses local linearization, controllability and observability Gramians, and Hankel singular values to analyze its internal computation. For a given input, we linearize the network around the corresponding hidden activation pattern and construct a state space model whose state consists of hidden neuron activations. The input state and state output Jacobians define local controllability and observability Gramians, from which we compute Hankel singular values and associated modes. These quantities provide a principled notion of neuron and pathway importance: controllability measures how easily each neuron can be excited by input perturbations, observability measures how strongly each neuron influences the output, and Hankel singular values rank internal modes that carry input output energy. We illustrate the framework on simple feedforward networks, including a 1 2 2 1 SwiGLU network and a 2 3 3 2 GELU network. By comparing different operating points, we show how activation saturation reduces controllability, shrinks the dominant Hankel singular value, and shifts the dominant internal mode to a different subset of neurons. The proposed method turns a neural network into a collection of local white box dynamical models and suggests which internal directions are natural candidates for pruning or constraints to improve interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。