arXiv:2504.08489math.STcs.LG2025-04

基于统计理论设计的深度学习方法,提升小样本回归性能。

Statistically guided deep learning

  • 结合优化、泛化与逼近的理论分析设计新算法
  • 在模拟数据上实现更优的有限样本误差表现
  • 适合关注理论指导实践的机器学习研究者

我们提出一种理论严谨的非参数回归深度学习算法。该算法使用具有逻辑激活函数的过参数化深层神经网络,通过梯度下降拟合数据。我们设计了特殊的网络拓扑结构、权重随机初始化方式,以及依赖数据的学习率和迭代步数选择策略。证明了该估计量的期望$L_2$误差的理论界,并在模拟数据上展示了其有限样本性能。结果表明,同时考虑优化、泛化与逼近的理论分析,可导出具有更好有限样本表现的新深度学习估计。

原文摘要 · Abstract (English)

We present a theoretically well-founded deep learning algorithm for nonparametric regression. It uses over-parametrized deep neural networks with logistic activation function, which are fitted to the given data via gradient descent. We propose a special topology of these networks, a special random initialization of the weights, and a data-dependent choice of the learning rate and the number of gradient descent steps. We prove a theoretical bound on the expected $L_2$ error of this estimate, and illustrate its finite sample size performance by applying it to simulated data. Our results show that a theoretical analysis of deep learning which takes into account simultaneously optimization, generalization and approximation can result in a new deep learning estimate which has an improved finite sample performance.

深度学习统计学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。