arXiv:2412.19677stat.MLcond-mat.dis-nn2024-12被引 3

揭示深度ReLU网络的唯一性恢复能力上限,发现4层即可接近最优扩张。

Deep ReLU networks -- injectivity capacity upper bounds

  • 基于随机对偶理论构建深层网络唯一性分析框架
  • 证明4层网络扩张趋近饱和,无需额外扩展即达最优性能
  • 为实际深度网络设计提供数学依据,适合理论与工程交叉研究者

本文研究深度ReLU前馈神经网络的唯一性恢复能力。针对任意给定的隐藏层结构,定义了其唯一性容量——即保证输入可由可行输出唯一恢复的最小输出与输入比。通过将单层网络的唯一性分析拓展至深层网络,建立了一个连接l层网络唯一性与ℓ₀球面感知机的l-扩展模型,从而大幅推广了文献[82]中关于单层唯一性与(1-扩展)ℓ₀球面感知机容量之间的同构关系。进一步提出基于随机对偶理论(RDT)的分析工具,用于统计处理扩展后的ℓ₀球面感知机性质,间接刻画深层ReLU网络特性。大量数值实验验证了该方法的实用性,并观测到显著的扩张饱和效应:仅需4层深度即可接近无需扩张的理想水平。这一结果与实际实验观察高度吻合,且此前未被任何现有数学方法解释。

原文摘要 · Abstract (English)

We study deep ReLU feed forward neural networks (NN) and their injectivity abilities. The main focus is on \emph{precisely} determining the so-called injectivity capacity. For any given hidden layers architecture, it is defined as the minimal ratio between number of network's outputs and inputs which ensures unique recoverability of the input from a realizable output. A strong recent progress in precisely studying single ReLU layer injectivity properties is here moved to a deep network level. In particular, we develop a program that connects deep $l$-layer net injectivity to an $l$-extension of the $\ell_0$ spherical perceptrons, thereby massively generalizing an isomorphism between studying single layer injectivity and the capacity of the so-called (1-extension) $\ell_0$ spherical perceptrons discussed in [82]. \emph{Random duality theory} (RDT) based machinery is then created and utilized to statistically handle properties of the extended $\ell_0$ spherical perceptrons and implicitly of the deep ReLU NNs. A sizeable set of numerical evaluations is conducted as well to put the entire RDT machinery in practical use. From these we observe a rapidly decreasing tendency in needed layers' expansions, i.e., we observe a rapid \emph{expansion saturation effect}. Only $4$ layers of depth are sufficient to closely approach level of no needed expansion -- a result that fairly closely resembles observations made in practical experiments and that has so far remained completely untouchable by any of the existing mathematical methodologies.

深度学习神经网络唯一性理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。