揭示深度神经网络分类器在高维数据下的最优收敛速率。
Optimal Convergence Rates of Deep Neural Network Classifiers
- 基于复合函数假设与Tsybakov噪声条件,推导出理论最优收敛率。
- 证明ReLU DNN用铰链损失可逼近该最优速率,仅差对数因子。
- 适用于高维分类任务,为实际应用提供理论支持。
本文研究在$[0,1]^d$上满足Tsybakov噪声条件(指数$ s \in [0,\infty] $)及复合假设的二分类问题。该假设要求条件概率函数为$ q+1 $个向量值多元函数的复合,每个分量函数为最大值函数或仅依赖于$d_*$个输入变量的Hölder-$β$光滑函数,其中$d_*$远小于输入维度$d$。我们证明,在此条件下,分类器超出0-1风险的最优收敛率为$\left( \frac{1}{n} \right)^{\frac{β\cdot(1\wedgeβ)^q}{{\frac{d_*}{s+1}+(1+\frac{1}{s+1})\cdotβ\cdot(1\wedgeβ)^q}}}$,且与输入维数$d$无关。此外,我们证明使用铰链损失训练的ReLU深度神经网络(DNN)可达到该最优收敛率,仅差一个对数因子。这一结果为ReLU DNN在高维分类任务中的优异表现提供了理论依据。该方法具有独立研究价值。
原文摘要 · Abstract (English)
In this paper, we study the binary classification problem on $[0,1]^d$ under the Tsybakov noise condition (with exponent $s \in [0,\infty]$) and the compositional assumption. This assumption requires the conditional class probability function of the data distribution to be the composition of $q+1$ vector-valued multivariate functions, where each component function is either a maximum value function or a Hölder-$β$ smooth function that depends only on $d_*$ of its input variables. Notably, $d_*$ can be significantly smaller than the input dimension $d$. We prove that, under these conditions, the optimal convergence rate for the excess 0-1 risk of classifiers is $\left( \frac{1}{n} \right)^{\frac{β\cdot(1\wedgeβ)^q}{{\frac{d_*}{s+1}+(1+\frac{1}{s+1})\cdotβ\cdot(1\wedgeβ)^q}}}$, which is independent of the input dimension $d$. Additionally, we demonstrate that ReLU deep neural networks (DNNs) trained with hinge loss can achieve this optimal convergence rate up to a logarithmic factor. This result provides theoretical justification for the excellent performance of ReLU DNNs in practical classification tasks, particularly in high-dimensional settings. The generalized approach is of independent interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。