研究卷积网络梯度流收敛性,证明在特定条件下可稳定到达临界点。
Convergence of gradient flow for learning convolutional neural networks
- 用梯度流分析线性卷积网络的学习过程
- 在训练数据满足弱条件时,梯度流必然收敛到临界点
- 为理解深度学习优化机制提供理论支持,适合研究者参考
卷积神经网络广泛应用于成像与图像识别。从训练数据中学习这类网络涉及非凸函数的最小化,使得标准优化方法(如随机梯度下降)的分析极具挑战。本文研究线性卷积网络的简化情形,证明当训练数据满足一定弱条件时,基于平方损失等损失函数定义的经验风险,其梯度流(可视为梯度下降的抽象)始终收敛至临界点。
原文摘要 · Abstract (English)
Convolutional neural networks are widely used in imaging and image recognition. Learning such networks from training data leads to the minimization of a non-convex function. This makes the analysis of standard optimization methods such as variants of (stochastic) gradient descent challenging. In this article we study the simplified setting of linear convolutional networks. We show that the gradient flow (to be interpreted as an abstraction of gradient descent) applied to the empirical risk defined via certain loss functions including the square loss always converges to a critical point, under a mild condition on the training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。