通过引入梯度信息,该方法让ReLU网络训练更快更准。
Sobolev acceleration for neural networks
- 在浅层网络中融合目标函数导数,改进优化路径
- 理论证明收敛速度提升,损失曲面条件数改善
- 适用于深度学习中的快速训练与泛化提升
Sobolev训练通过将目标函数的导数融入损失函数,在对比传统L²训练时表现出加速收敛和更好泛化能力。然而其内在机制仍不完全清楚。本文首次建立严格理论框架,证明在高斯输入与浅层学生-教师结构下,对ReLU网络的Sobolev训练能加速收敛。我们推导出种群梯度与海森矩阵的精确公式,并量化了损失景观条件性改善及梯度流收敛速率提升。大量数值实验验证了理论结果,表明该优势可扩展至现代深度学习任务。
原文摘要 · Abstract (English)
Sobolev training, which integrates target derivatives into the loss functions, has been shown to accelerate convergence and improve generalization compared to conventional $L^2$ training. However, the underlying mechanisms of this training method remain only partially understood. In this work, we present the first rigorous theoretical framework proving that Sobolev training accelerates the convergence of Rectified Linear Unit (ReLU) networks. Under a student-teacher framework with Gaussian inputs and shallow architectures, we derive exact formulas for population gradients and Hessians, and quantify the improvements in conditioning of the loss landscape and gradient-flow convergence rates. Extensive numerical experiments validate our theoretical findings and show that the benefits of Sobolev training extend to modern deep learning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。