arXiv:2602.19691stat.MLcs.LG2026-02被引 1

平滑激活函数让浅层网络自适应光滑性,实现最优逼近与学习率。

Smoothness Adaptivity in Constant-Depth Neural Networks: Optimal Rates via Smooth Activations

  • 用多尺度逼近框架设计平滑激活的浅层网络
  • 宽度增加即可达最优误差率(对数因子内)
  • 适合研究理论深度学习或优化模型泛化者

平滑激活函数在现代深度学习中普遍使用,但其相对于非平滑激活的理论优势尚不清晰。本文研究了具有平滑激活的常深神经网络在 Sobolev 空间 $W^{s, ty}([0,1]^d)$($s>0$)上学习函数的逼近与统计性质。证明了:仅通过增加宽度,常深网络即可实现光滑性自适应,达到最小最大最优逼近与估计误差率(对数因子内)。而对非平滑激活如 ReLU,光滑性自适应受深度根本限制:逼近阶受限于深度,更高阶光滑性需成比例增加深度。结果表明,激活平滑性是与深度互补的关键机制,可实现对 Sobolev 函数类的最优率。技术上,分析基于多尺度逼近框架,构造出参数范数与模型规模可控的显式神经网络逼近器。该复杂度控制确保了经验风险最小化下的统计可学习性,避免了以往分析中常见的 $ ilde{ ext{O}}( rac{d}{ ext{poly}(n)})$-sparsity 约束。

原文摘要 · Abstract (English)

Smooth activation functions are ubiquitous in modern deep learning, yet their theoretical advantages over non-smooth counterparts remain poorly understood. In this work, we study both approximation and statistical properties of neural networks with smooth activations for learning functions in the Sobolev space $W^{s,\infty}([0,1]^d)$ with $s>0$. We prove that constant-depth networks equipped with smooth activations achieve smoothness adaptivity: increasing width alone suffices to attain the minimax-optimal approximation and estimation error rates (up to logarithmic factors). In contrast, for non-smooth activations such as ReLU, smoothness adaptivity is fundamentally limited by depth: the attainable approximation order is bounded by depth, and higher-order smoothness requires proportional depth growth. These results identify activation smoothness as a fundamental mechanism, complementary to depth, for achieving optimal rates over Sobolev function classes. Technically, our analysis is based on a multi-scale approximation framework that yields explicit neural network approximators with controlled parameter norms and model size. This complexity control ensures statistical learnability under empirical risk minimization (ERM) and avoids the impractical $\ell^0$-sparsity constraints commonly required in prior analyses.

神经网络平滑激活逼近理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。