arXiv:2505.07792stat.MLcond-mat.dis-nn2025-05被引 3

解析了两层网络中dropout的理论机制,揭示其降噪与正则化原理。

Analytic theory of dropout regularization

  • 基于高维极限推导出描述训练过程的微分方程,解析dropout动态
  • 发现最优dropout率随数据噪声升高而增加,且能抑制隐藏单元间有害相关性
  • 结果被大量模拟验证,适用于理解与优化深度学习正则化策略

Dropout是一种广泛用于人工神经网络训练的正则化技术,通过在训练过程中动态关闭部分网络单元来提升模型鲁棒性。尽管应用广泛,其概率选择多依赖经验,理论解释仍不充分。本文针对在线随机梯度下降训练的两层神经网络,在高维极限下,推导出完整刻画网络演化过程的一组常微分方程,准确捕捉dropout的影响。我们获得了关于泛化误差和不同训练阶段最优dropout概率的精确结果。分析表明,dropout能有效减少隐藏节点间的有害相关性,降低标签噪声的影响,且最优dropout概率随数据噪声水平升高而增大。研究结果经大量数值模拟验证。

原文摘要 · Abstract (English)

Dropout is a regularization technique widely used in training artificial neural networks to mitigate overfitting. It consists of dynamically deactivating subsets of the network during training to promote more robust representations. Despite its widespread adoption, dropout probabilities are often selected heuristically, and theoretical explanations of its success remain sparse. Here, we analytically study dropout in two-layer neural networks trained with online stochastic gradient descent. In the high-dimensional limit, we derive a set of ordinary differential equations that fully characterize the evolution of the network during training and capture the effects of dropout. We obtain a number of exact results describing the generalization error and the optimal dropout probability at short, intermediate, and long training times. Our analysis shows that dropout reduces detrimental correlations between hidden nodes, mitigates the impact of label noise, and that the optimal dropout probability increases with the level of noise in the data. Our results are validated by extensive numerical simulations.

dropout正则化理论分析神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。