平滑激活函数可缓解深度网络均匀收敛的维数灾难问题
Mitigating the Curse of Dimensionality in Uniform Convergence of Deep Neural Networks via Smooth Activations

- 用平滑激活函数构建深层网络,利用目标函数的低维分层结构
- 在多种回归任务中实现非渐近均匀收敛,避免维数灾难
- 适合需要最坏情况可靠性保障的统计学习任务
本文建立了光滑激活深度神经网络(smooth DNN)统一收敛性的理论框架。尽管标准ReLU网络在各类非参数回归任务中于$L^2(P)$范数下达到极小极大最优率,但本文证明最小二乘ReLU估计器在统一收敛性上仍可能受维数灾难影响。为满足下游任务对最坏情况可靠性的要求,我们分析了包含前馈与残差结构的光滑激活DNN。建立了新的伪维数界、非渐近逼近保证及霍尔德范数界,进而推导出光滑DNN在Huber、最小二乘、分位数和逻辑回归等多种统计场景下的非渐近统一收敛速率。证明光滑DNN可通过自适应利用目标函数的低维分层组合结构,缓解统一收敛中的维数灾难。仿真研究与真实应用验证了结果的有效性,表明光滑DNN是具有理论依据且实际可行的替代方案。
原文摘要 · Abstract (English)
This paper establishes a theoretical framework for the uniform convergence of smoothly activated deep neural network (DNN) estimators. While standard ReLU networks achieve minimax-optimal rates in the $L^2(P)$ norm for various nonparametric regression tasks, we establish a theoretical lower bound demonstrating that least-squares ReLU estimators can suffer from the curse of dimensionality in their uniform convergence behavior. Motivated by the need for reliable uniform guarantees in downstream tasks requiring worst-case reliability, we address this limitation by analyzing smoothly activated DNNs (smooth DNNs), encompassing both feedforward and residual structures. We establish novel pseudo-dimension bounds, non-asymptotic approximation guarantees, and Hölder-norm bounds for the approximators of these models. Leveraging these results, we derive non-asymptotic uniform convergence rates for smooth DNN estimators across multiple statistical contexts, including Huber, least-squares, quantile, and logistic regression. We prove that smooth DNNs can mitigate the {curse of dimensionality} in uniform convergence by adaptively exploiting the low-dimensional hierarchical composition structure of the target function. Supported by both simulation studies and a real-world application, our results position smooth DNNs as a theoretically grounded and practically viable alternative to ReLU networks for statistical learning tasks requiring uniform guarantees.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。