深度对解析函数逼近比宽度更重要,揭示了深层网络的潜力。
Approximation of Analytic Functions by ReLU Neural Networks with Adjustable Depth and Width
- 用宽度和深度联合刻画网络性能,突破传统单一参数限制。
- 发现解析函数逼近率可达指数级衰减:$\mathcal{O}(N^{-CL^\tau})$。
- 适用于需要高精度逼近的数学建模与理论分析场景。
与多数仅以总参数量为单一指标的研究不同,本文采用宽度 $N$ 和深度 $L$ 联合刻画神经网络逼近能力,赋予架构更大灵活性。现有工作针对光滑度有限的函数类,得到典型逼近率为 $\mathcal{O}(N^{-2s/d}L^{-2s/d})$,表明深度与宽度作用对称。本文则聚焦光滑度无限的解析函数,建立基于 $(N,L)$-刻画的上界:逼近率可达 $\mathcal{O}(N^{-CL^\tau})$,其中 $C>0$ 为常数,$\tau>0$ 受 $L$ 与 $N$ 关系影响;当 $N$ 约为 $L^d$ 时,$\tau=1$。结果表明,在解析函数逼近中,深度比宽度更具决定性作用。关键技术难点在于平滑参数与逼近精度间的权衡,为此我们精心构造了用于逼近幂函数、多元乘积和多项式的若干 ReLU 网络,其方法本身或具独立价值。
原文摘要 · Abstract (English)
In contrast to most studies on neural network approximation theory that characterize results through a single parameter, such as the total number of network parameters, \cite{shen2020deep} pioneered the characterization of approximation rates as a joint function of the width parameter $N$ and the depth parameter $L$, thereby granting greater architectural flexibility. Existing works using the $(N,L)$-characterization focus on function classes with finite smoothness $s$, establishing a typical approximation rate of $\mathcal{O}\left(N^{-2s/d}L^{-2s/d}\right)$ with $d$ denoting the input dimension, which indicates that network depth and width play symmetric roles for these classes. In contrast, this paper establishes upper bounds for the approximation of analytic functions, which possess infinite smoothness, via ReLU networks under the $(N,L)$-characterization. Specifically, we derive approximation rates of $\mathcal{O}\left(N^{-C L^τ}\right)$, where $C>0$ is some constant and $τ>0$ is a parameter influenced by the relation between $L$ and $N$. In particular, $τ=1$ if $N$ scales roughly as $L^d$. Our findings reveal that depth plays a more critical role than width in the context of analytic function approximation. The main technical difficulty of obtaining such upper bounds lies in the trade-off between the smoothness parameters and the approximation accuracy. To overcome this difficulty, we employ refined constructions of several ReLU networks to approximate power functions, multivariate multiplication, and polynomials, which may be of independent interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。