arXiv:2410.21702stat.MLcs.LG2024-10被引 5

证明深度神经网络在依赖数据下仍具最优估计性能

Minimax optimality of deep neural networks on dependent data via PAC-Bayes bounds

  • 放宽独立同分布假设,允许马尔可夫依赖数据
  • 推导广义贝叶斯估计器风险上界,匹配已有下界
  • 适用于回归与分类,适合理论研究者参考

Schmidt-Hieber(2020)首次证明了带ReLU激活函数的深度神经网络在一大类复合函数上的最小最大最优性。本文将其拓展至多个方向:首先,去除观测值独立同分布假设,允许存在时间依赖性,假设观测为具有非零伪谱间隙的马尔可夫链;其次,研究更广泛的机器学习问题,包括最小二乘与逻辑回归作为特例。借助PAC-Bayes奥数不等式及Paulin(2015)的伯恩斯坦不等式,推导出广义贝叶斯估计器的估计风险上界。在最小二乘回归情形,该上界与Schmidt-Hieber(2020)的下界仅差对数因子。针对逻辑损失分类问题,建立类似下界,并证明所提DNN估计器在最小最大意义下最优。

原文摘要 · Abstract (English)

In a groundbreaking work, Schmidt-Hieber (2020) proved the minimax optimality of deep neural networks with ReLu activation for least-square regression estimation over a large class of functions defined by composition. In this paper, we extend these results in many directions. First, we remove the i.i.d. assumption on the observations, to allow some time dependence. The observations are assumed to be a Markov chain with a non-null pseudo-spectral gap. Then, we study a more general class of machine learning problems, which includes least-square and logistic regression as special cases. Leveraging on PAC-Bayes oracle inequalities and a version of Bernstein inequality due to Paulin (2015), we derive upper bounds on the estimation risk for a generalized Bayesian estimator. In the case of least-square regression, this bound matches (up to a logarithmic factor) the lower bound of Schmidt-Hieber (2020). We establish a similar lower bound for classification with the logistic loss, and prove that the proposed DNN estimator is optimal in the minimax sense.

深度学习统计学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。