用图论解析ReLU网络结构,揭示其表达与泛化关系
The Geometry of ReLU Networks through the ReLU Transition Graph
- 构建激活模式变化的神经元翻转图(RTG),刻画网络线性区域
- 证明RTG连通性并给出大小与直径的紧致界,关联熵与平均度到泛化误差
- 为模型压缩、正则化提供新思路,适合研究网络结构与泛化者
我们提出一种新的理论框架,通过称为ReLU转换图(RTG)的组合对象分析ReLU神经网络。图中每个节点对应网络激活模式产生的一个线性区域,边连接仅通过单个神经元翻转即可相互转换的区域。基于该结构,我们推导出一系列新理论结果,将RTG的几何性质与表达能力、泛化性能及鲁棒性联系起来。主要贡献包括:对RTG大小和直径的紧致组合上界、证明了RTG的连通性,以及对VC维的图论解释。此外,我们将熵和平均度与泛化误差相关联。所有理论结果均通过在不同网络深度、宽度和数据设置下的精心控制实验得到验证。本工作首次以图论统一处理ReLU网络结构,为基于RTG分析的压缩、正则化和复杂度控制开辟新路径。
原文摘要 · Abstract (English)
We develop a novel theoretical framework for analyzing ReLU neural networks through the lens of a combinatorial object we term the ReLU Transition Graph (RTG). In this graph, each node corresponds to a linear region induced by the network's activation patterns, and edges connect regions that differ by a single neuron flip. Building on this structure, we derive a suite of new theoretical results connecting RTG geometry to expressivity, generalization, and robustness. Our contributions include tight combinatorial bounds on RTG size and diameter, a proof of RTG connectivity, and graph-theoretic interpretations of VC-dimension. We also relate entropy and average degree of the RTG to generalization error. Each theoretical result is rigorously validated via carefully controlled experiments across varied network depths, widths, and data regimes. This work provides the first unified treatment of ReLU network structure via graph theory and opens new avenues for compression, regularization, and complexity control rooted in RTG analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。