解析神经网络逼近能力的理论,揭示深度与宽度如何影响学习效率。
Approximation Theory for Neural Networks: Old and New
- 从经典密度定理到量化误差分析,建立网络规模与函数平滑性的关系。
- 证明深层网络在结构化函数上比浅层网络更节省参数。
- 对比传统前馈网络与新型KAN架构的理论优势,适合研究者参考。
通用逼近定理为神经网络的表达能力提供了数学解释。它们表明,在激活函数满足一定温和条件时,前馈神经网络在连续函数、L^p空间或Sobolev空间等广泛函数类中是稠密的。过去四十年间,这些定性通用性结果已发展成丰富的定量理论,涵盖逼近速率、参数效率以及深度与宽度等结构特征的作用。本文综述了该理论的若干方面:回顾单隐层网络的经典稠密性结果,以及逼近误差与网络规模和目标函数光滑性之间的量化界限。特别关注深度与宽度的权衡,以及深层结构在结构化函数类中可实现更高参数效率的成果。此外,还介绍了最近关于柯尔莫戈洛夫-阿诺德网络(KANs)的研究进展,其作为替代性架构,其逼近理论性质正获得越来越多的理论关注。
原文摘要 · Abstract (English)
Universal approximation theorems provide a mathematical explanation for the expressive power of neural networks. They assert that, under mild conditions on the activation function, feedforward neural networks are dense in broad function classes, such as continuous functions on compact subsets of $\mathbb{R}^d$, $L^p$ spaces, or Sobolev spaces. Over the past four decades, these qualitative universality results have evolved into a rich quantitative theory addressing approximation rates, parameter efficiency, and the role of architectural features such as depth and width. This survey presents several glimpses into this theory. We review classical density results for single-hidden-layer networks, as well as quantitative bounds that relate approximation error to network size and smoothness assumptions on target functions. Particular emphasis is placed on depth--width trade-offs and on results demonstrating that deeper architectures can achieve superior parameter efficiency for structured function classes. In addition to standard feedforward neural networks, we also review recent developments on Kolmogorov--Arnold Networks (KANs), which offer an alternative architectural paradigm and whose approximation-theoretic properties have begun to attract significant theoretical attention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。