揭示神经网络训练中鲁棒性动态变化的数学规律
Optimization-Induced Dynamics of Lipschitz Continuity in Neural Networks
- 用随机微分方程建模训练时的 Lipschitz 连续性演变
- 发现梯度流、噪声投影与海森矩阵共同驱动变化
- 适用于关注模型鲁棒性与训练机制的研究者
Lipschitz 连续性刻画神经网络对输入微小扰动的最坏敏感度;然而其在训练过程中的动态演化仍缺乏深入研究。本文提出一个严格的数学框架,通过随机微分方程(SDEs)建模随机梯度下降(SGD)训练过程中 Lipschitz 连续性的时变特性,同时捕捉确定性与随机性作用力。理论分析揭示三个主要驱动力:(i) 优化动力学诱导的梯度流在参数矩阵算子范数雅可比矩阵上的投影;(ii) 小批量采样随机性带来的梯度噪声在算子范数雅可比矩阵上的投影;(iii) 梯度噪声在参数矩阵算子范数海森矩阵上的投影。此外,该框架还阐明了噪声监督、参数初始化、批次大小及小批量采样轨迹等因素如何影响 Lipschitz 连续性的演化。实验结果表明理论推论与观测行为高度一致。
原文摘要 · Abstract (English)
Lipschitz continuity characterizes the worst-case sensitivity of neural networks to small input perturbations; yet its dynamics (i.e. temporal evolution) during training remains under-explored. We present a rigorous mathematical framework to model the temporal evolution of Lipschitz continuity during training with stochastic gradient descent (SGD). This framework leverages a system of stochastic differential equations (SDEs) to capture both deterministic and stochastic forces. Our theoretical analysis identifies three principal factors driving the evolution: (i) the projection of gradient flows, induced by the optimization dynamics, onto the operator-norm Jacobian of parameter matrices; (ii) the projection of gradient noise, arising from the randomness in mini-batch sampling, onto the operator-norm Jacobian; and (iii) the projection of the gradient noise onto the operator-norm Hessian of parameter matrices. Furthermore, our theoretical framework sheds light on such as how noisy supervision, parameter initialization, batch size, and mini-batch sampling trajectories, among other factors, shape the evolution of the Lipschitz continuity of neural networks. Our experimental results demonstrate strong agreement between the theoretical implications and the observed behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。