用混沌瞬态加速神经网络训练,突破传统梯度下降局限
Leveraging chaotic transients in the training of artificial neural networks
- 在大学习率下引入混沌瞬态,实现探索与利用的平衡
- 测试集准确率达最优时,训练时间最短,存在最佳混沌区
- 适用于多种模型与任务,为高效训练提供新思路
传统神经网络优化算法多为梯度下降等依赖型松弛动力学。本文研究了大学习率下的训练轨迹动态特性,发现当学习率处于某一区间时,优化过程从纯利用型转向探索-利用平衡状态:网络仍可学习,但轨迹对初始条件敏感(表现为正的最大李雅普诺夫指数)。有趣的是,达到可接受测试精度所需的训练时间在此区域达到最小值,表明可通过定位混沌临界点加速训练。该现象最初在MNIST分类任务中验证,且在多种学习架构(包括浅层与深层全连接网络、卷积神经网络)和超参数(不同激活函数、权重正则化)下均具定性一致性,揭示了瞬态混沌动力学在神经网络训练中的涌现性构造作用。
原文摘要 · Abstract (English)
Traditional algorithms to optimize artificial neural networks when confronted with a supervised learning task are usually exploitation-type relaxational dynamics such as gradient descent (GD). Here, we explore the dynamics of the neural network trajectory along training for unconventionally large learning rates. We show that for a region of values of the learning rate, the GD optimization shifts away from purely exploitation-like algorithm into a regime of exploration-exploitation balance, as the neural network is still capable of learning but the trajectory shows sensitive dependence on initial conditions --as characterized by positive network maximum Lyapunov exponent--. Interestingly, the characteristic training time required to reach an acceptable accuracy in the test set reaches a minimum precisely in such learning rate region, further suggesting that one can accelerate the training of artificial neural networks by locating at the onset of chaos. Our results --initially illustrated for the MNIST classification task-- qualitatively hold for a range of supervised learning tasks, {learning architectures (including both shallow and deep multilayer perceptrons and convolutional neural networks) and other hyperparameters (different activation functions and weight regularisation),} and showcase the emergent, constructive role of transient chaotic dynamics in the training of artificial neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。