用动力系统视角解析神经网络的训练与信息传播机制
A Dynamical Systems Perspective on the Analysis of Neural Networks
- 将神经网络训练和信息传递转化为动态系统问题
- 揭示了过参数化下的稳定性边界现象及隐式偏差机制
- 适用于理解生成模型与梯度训练中的根本难题
本文从动力系统角度分析机器学习算法的多个方面。我们将深度神经网络、(随机)梯度下降等众多挑战重新表述为动态命题。首先,研究神经网络中信息的传播过程,解释了增强型神经微分方程对给定正则性函数的通用嵌入性质,分类了多层感知机与神经微分方程所表征的函数类,并分析了神经延迟方程中的记忆依赖性。其次,从动力系统视角分析训练过程,描述梯度下降的稳定性,研究过定问题下的稳定性,并拓展至过参数化情形,揭示“稳定边界”现象及其对隐式偏差的可能解释。针对随机梯度下降,在过参数化设置下通过插值解的李雅普诺夫指数给出稳定性结果。第三,阐述神经网络的均场极限相关成果,提出将异质神经网络纳入图极限框架的新方法,表明大规模神经网络自然属于图上的库朗托型模型及其大图极限。最后指出,类似策略可用于可解释性和可靠性人工智能、生成模型,以及梯度训练中的基本问题如反向传播、梯度消失/爆炸等。
原文摘要 · Abstract (English)
In this chapter, we utilize dynamical systems to analyze several aspects of machine learning algorithms. As an expository contribution we demonstrate how to re-formulate a wide variety of challenges from deep neural networks, (stochastic) gradient descent, and related topics into dynamical statements. We also tackle three concrete challenges. First, we consider the process of information propagation through a neural network, i.e., we study the input-output map for different architectures. We explain the universal embedding property for augmented neural ODEs representing arbitrary functions of given regularity, the classification of multilayer perceptrons and neural ODEs in terms of suitable function classes, and the memory-dependence in neural delay equations. Second, we consider the training aspect of neural networks dynamically. We describe a dynamical systems perspective on gradient descent and study stability for overdetermined problems. We then extend this analysis to the overparameterized setting and describe the edge of stability phenomenon, also in the context of possible explanations for implicit bias. For stochastic gradient descent, we present stability results for the overparameterized setting via Lyapunov exponents of interpolation solutions. Third, we explain several results regarding mean-field limits of neural networks. We describe a result that extends existing techniques to heterogeneous neural networks involving graph limits via digraph measures. This shows how large classes of neural networks naturally fall within the framework of Kuramoto-type models on graphs and their large-graph limits. Finally, we point out that similar strategies to use dynamics to study explainable and reliable AI can also be applied to settings such as generative models or fundamental issues in gradient training methods, such as backpropagation or vanishing/exploding gradients.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。