通过路径约束训练,让ReLU网络权重更稳定高效地收敛。
Path-conditioned training: a principled way to rescale ReLU neural networks
- 基于路径提升框架,设计几何准则重标定网络参数。
- 实验显示该方法可显著加速随机初始化网络的训练过程。
- 适合关注网络初始化与训练动力学的深度学习研究者。
尽管近期算法取得进展,我们仍缺乏系统性方法来利用ReLU神经网络参数中已知的缩放对称性。虽然经过适当缩放的权重实现相同函数,但训练动态可能截然不同。本文基于最新的路径提升框架,提出一种几何动机的重标定准则,其最小化过程可导出一种对齐策略,使路径提升空间中的核与指定参考对齐。我们推导出高效的对齐算法。在随机初始化背景下,分析了网络结构与初始化尺度如何共同影响结果。数值实验表明该方法具有加速训练的潜力。
原文摘要 · Abstract (English)
Despite recent algorithmic advances, we still lack principled ways to leverage the well-documented rescaling symmetries in ReLU neural network parameters. While two properly rescaled weights implement the same function, the training dynamics can be dramatically different. To offer a fresh perspective on exploiting this phenomenon, we build on the recent path-lifting framework, which provides a compact factorization of ReLU networks. We introduce a geometrically motivated criterion to rescale neural network parameters which minimization leads to a conditioning strategy that aligns a kernel in the path-lifting space with a chosen reference. We derive an efficient algorithm to perform this alignment. In the context of random network initialization, we analyze how the architecture and the initialization scale jointly impact the output of the proposed method. Numerical experiments illustrate its potential to speed up training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。