arXiv:2509.03677cs.LGcs.AI2025-09

提出无需调参的梯度归一化方法,稳定训练并提升模型泛化能力。

Insights from Gradient Dynamics: Gradient Autoscaled Normalization

  • 根据梯度动态演化规律设计自适应归一化策略
  • 在CIFAR-100上对多个模型提升或保持测试准确率
  • 为优化研究提供可追踪梯度行为的新视角

梯度动态在决定深度神经网络的稳定性与泛化能力中起核心作用。本文通过实验分析卷积网络中梯度方差与标准差随训练过程的演变,发现各层及全局尺度均呈现一致变化规律。基于此,我们提出一种无超参数的梯度归一化方法,使梯度缩放与其自然演化对齐,防止意外放大,稳定优化过程,并保留收敛性保证。在包含ResNet-20、ResNet-56和VGG-16-BN的CIFAR-100挑战性基准上的实验表明,该方法在强泛化条件下仍能维持或提升测试精度。本研究不仅提升实际性能,更强调直接追踪梯度动态的重要性,旨在弥合理论预期与实际行为之间的差距,为未来优化研究提供洞见。

原文摘要 · Abstract (English)

Gradient dynamics play a central role in determining the stability and generalization of deep neural networks. In this work, we provide an empirical analysis of how variance and standard deviation of gradients evolve during training, showing consistent changes across layers and at the global scale in convolutional networks. Motivated by these observations, we propose a hyperparameter-free gradient normalization method that aligns gradient scaling with their natural evolution. This approach prevents unintended amplification, stabilizes optimization, and preserves convergence guarantees. Experiments on the challenging CIFAR-100 benchmark with ResNet-20, ResNet-56, and VGG-16-BN demonstrate that our method maintains or improves test accuracy even under strong generalization. Beyond practical performance, our study highlights the importance of directly tracking gradient dynamics, aiming to bridge the gap between theoretical expectations and empirical behaviors, and to provide insights for future optimization research.

梯度分析优化器泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。