arXiv:2409.08935cs.LGcs.AI2024-09被引 2

首次为权重归一化网络提供优化与泛化理论保障。

Optimization and Generalization Guarantees for Weight Normalization

  • 通过分析损失函数的海森矩阵,建立权重归一化网络的优化收敛性。
  • 证明泛化误差与网络深度呈次线性关系,与宽度无关。
  • 实验验证理论中的归一化项与训练性能的关联性。

权重归一化(WeightNorm)在深度神经网络训练中广泛应用,现代深度学习库也内置了其实现。本文首次对具有光滑激活函数的深层权重归一化模型的优化与泛化性质进行了理论刻画。针对优化问题,从损失函数的海森矩阵形式出发,发现预测器的海森矩阵较小时可实现可处理分析,因此我们界定了权重归一化网络海森矩阵谱范数,并揭示其对网络宽度和归一化项的依赖关系——后者是无归一化网络独有的特性。在此基础上,在合适假设下建立了梯度下降的训练收敛性保证。针对泛化问题,利用权重归一化获得基于一致收敛的泛化界,该界与网络宽度无关,且关于深度呈次线性依赖。最后,我们提供了实验结果,展示归一化项及其他理论相关量如何影响权重归一化网络的训练过程。

原文摘要 · Abstract (English)

Weight normalization (WeightNorm) is widely used in practice for the training of deep neural networks and modern deep learning libraries have built-in implementations of it. In this paper, we provide the first theoretical characterizations of both optimization and generalization of deep WeightNorm models with smooth activation functions. For optimization, from the form of the Hessian of the loss, we note that a small Hessian of the predictor leads to a tractable analysis. Thus, we bound the spectral norm of the Hessian of WeightNorm networks and show its dependence on the network width and weight normalization terms--the latter being unique to networks without WeightNorm. Then, we use this bound to establish training convergence guarantees under suitable assumptions for gradient decent. For generalization, we use WeightNorm to get a uniform convergence based generalization bound, which is independent from the width and depends sublinearly on the depth. Finally, we present experimental results which illustrate how the normalization terms and other quantities of theoretical interest relate to the training of WeightNorm networks.

权重归一化优化理论泛化界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。