将深度神经网络训练动态分解为内在结构与数据交互部分,揭示其物理规律。
Bulk-boundary decomposition of neural networks
- 从随机梯度下降出发,分离出与数据无关的体部项和依赖数据的边界项。
- 发现网络内部存在能量连续性方程,体现局部与均匀性本质。
- 适合研究模型训练机制的理论学者,理解深层网络动态的新视角。
我们提出一种新的框架——体-边界分解,用于理解深度神经网络的训练动态。基于随机梯度下降形式,我们证明拉格朗日量可重新组织为一个与数据无关的体部项和一个依赖数据的边界项。体部项捕捉由网络架构和激活函数决定的内在动力学,边界项则反映输入层与输出层训练样本带来的随机相互作用。该分解揭示了深层网络底层的局部性和同质性结构。作为局部性与同质性的物理后果,我们在深层神经网络中推导出能量连续性方程。
原文摘要 · Abstract (English)
We present the bulk--boundary decomposition as a new framework for understanding the training dynamics of deep neural networks. Starting from the stochastic gradient descent formulation, we show that the Lagrangian can be reorganized into a data-independent bulk term and a data-dependent boundary term. The bulk captures the intrinsic dynamics set by network architecture and activation functions, while the boundary reflects stochastic interactions from training samples at the input and output layers. This decomposition exposes the local and homogeneous structure underlying deep networks. As a physical consequence of locality and homogeneity, we derive the energy continuity equation within a deep neural network.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。