提出一种新方法,精准初始化门控循环网络的权重以提升性能。
A Random-Matrix Criterion for Initializing Gated Recurrent Neural Networks

- 基于随机矩阵理论,推导出适用于多种循环架构的临界初始化准则。
- 在混沌预测任务中,该准则对应的权重方差使门控网络性能达到峰值。
- 为未来神经网络初始化设计提供可推广的理论指导,适合模型开发者参考。
训练前的权重初始化一直是推动深度学习发展的关键因素之一。在‘储备池计算’中,读出层权重通过线性学习,而储备池权重固定,极大影响系统动态的丰富性、稳定性和记忆能力。在无限宽度极限下,有意义的初始化需位于随机初始化模型的有效临界点上。这一相变由权重方差 $g^2$ 控制,将有序相与信息逐渐退化的混沌相分隔开。本文推导出一种简单准则,用于估算一大类循环架构的临界值 $g_c$,并证明该准则能紧密追踪门控循环网络在混沌预测任务中的最佳性能增益点。最后,我们主张该准则可作为未来初始化方案的设计原则。
原文摘要 · Abstract (English)
Proper weight initialization prior to training has historically been one of the key factors that helped kick off the deep learning revolution. Initialization is even more crucial in "reservoir computing", where the weights of a readout layer are learned linearly while the reservoir weights are fixed and largely determine the richness, stability and memory of the resulting dynamics. In the infinite-width limit it has been shown that meaningful initializations are those sitting at an effective critical point of the randomly initialized model. The phase transition is controlled by the weight variance $g^2$ and separates an ordered phase from a chaotic one where information progressively degrades. Here we derive a simple criterion to estimate the critical $g_c$ for a broad class of recurrent architectures and we show that it closely tracks the gain at which a gated-RNN reservoir achieves peak performance on a chaotic forecasting task. Finally, we argue that our criterion can serve as a design principle for future initialization schemes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。