arXiv:2504.07371cs.LG2025-04ICML被引 3

揭示了通用逼近所需最小网络宽度,适用于多种激活函数。

Minimum width for universal approximation using squashable activation functions

  • 提出可压缩激活函数概念,统一分析网络宽度需求。
  • 证明最小宽度为 max{dx, dy, 2},除非 dx=dy=1 且非单调。
  • 涵盖非仿射解析函数与分段函数,适用范围广。

已知仅对 ReLU 及其变体确定了无界深度网络的精确最小宽度。本文研究一般激活函数下的最小宽度问题,聚焦于可压缩函数——即可通过交替复合仿射变换逼近恒等函数和二值阶跃函数的激活函数。我们证明:使用可压缩激活函数的网络若要实现从 [0,1]^{d_x} 到 ℝ^{d_y} 的 L^p 函数的普遍逼近,其最小宽度为 max{d_x, d_y, 2},除非 d_x = d_y = 1;当 d_x = d_y = 1 时,若激活函数单调,该界仍成立。我们进一步给出可压缩性的充分条件,证明所有非仿射解析函数及一类分段函数均满足可压缩性,因此该最小宽度结果适用于这些广泛类别的激活函数。

原文摘要 · Abstract (English)

The exact minimum width that allows for universal approximation of unbounded-depth networks is known only for ReLU and its variants. In this work, we study the minimum width of networks using general activation functions. Specifically, we focus on squashable functions that can approximate the identity function and binary step function by alternatively composing with affine transformations. We show that for networks using a squashable activation function to universally approximate $L^p$ functions from $[0,1]^{d_x}$ to $\mathbb R^{d_y}$, the minimum width is $\max\{d_x,d_y,2\}$ unless $d_x=d_y=1$; the same bound holds for $d_x=d_y=1$ if the activation function is monotone. We then provide sufficient conditions for squashability and show that all non-affine analytic functions and a class of piecewise functions are squashable, i.e., our minimum width result holds for those general classes of activation functions.

神经网络宽度下限激活函数通用逼近

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。