从图论角度分析神经网络,揭示其保持信息完整性的条件。
A Graph Sufficiency Perspective for Neural Networks
- 将神经网络层视为基于图的变换,通过锚点和输入间的配对函数建模。
- 在无限宽或分区域输入分布下,证明层输出可保留输入的条件分布信息。
- 适用于全连接、卷积网络,为深度学习提供新的统计解释视角。
本文通过图变量与统计充分性分析神经网络。将神经网络层视为基于图的变换,其中神经元作为输入与学习锚点之间的成对函数。在此框架下,我们建立了层输出对层输入充分的条件,即每一层均保持目标变量关于输入变量的条件分布。本文探索了两种理论路径:第一条假设密集锚点,在无限宽度极限下证明渐近充分性且训练中保持不变;第二条更贴近实际架构,通过假设输入分布区域分离并构建合适的锚点,证明有限宽度网络中的精确或近似充分性。该路径可保证无限层数下的充分性,并为标准神经网络在回归与分类任务中的最优损失提供误差界。本框架涵盖全连接层、一般成对函数、ReLU与Sigmoid激活以及卷积神经网络。整体工作将统计充分性、图论表示与深度学习相融合,为神经网络提供了新的统计理解。
原文摘要 · Abstract (English)
This paper analyzes neural networks through graph variables and statistical sufficiency. We interpret neural network layers as graph-based transformations, where neurons act as pairwise functions between inputs and learned anchor points. Within this formulation, we establish conditions under which layer outputs are sufficient for the layer inputs, that is, each layer preserves the conditional distribution of the target variable given the input variable. We explore two theoretical paths under this graph-based view. The first path assumes dense anchor points and shows that asymptotic sufficiency holds in the infinite-width limit and is preserved throughout training. The second path, more aligned with practical architectures, proves exact or approximate sufficiency in finite-width networks by assuming region-separated input distributions and constructing appropriate anchor points. This path can ensure the sufficiency property for an infinite number of layers, and provide error bounds on the optimal loss for both regression and classification tasks using standard neural networks. Our framework covers fully connected layers, general pairwise functions, ReLU and sigmoid activations, and convolutional neural networks. Overall, this work bridges statistical sufficiency, graph-theoretic representations, and deep learning, providing a new statistical understanding of neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。