提出新型激活更新方法,突破传统归一化限制,提升模型优化效果。
The Affine Divergence: Aligning Activation Updates Beyond Normalisation
- 基于激活梯度的非理想缩放问题,设计新更新机制
- 新方法在多个测试中优于传统归一化,无需尺度不变性
- 适用于卷积与注意力层,适合研究优化机制的学者
梯度下降中,参数更新虽沿最陡下降方向,但激活更新却未达到最优。激活更接近损失函数且携带样本依赖信息,但其在仿射、卷积及注意力层中存在非理想的样本级缩放。现有修正方法虽简单,却能从原理推导出归一化,揭示了归一化的全新机理。进一步提出一种功能上不同于现代归一化的替代方案,无需尺度不变性却表现优异。该方法推广至卷积层形成新形式'PatchNorm',为不可分解的归一化器。整体构成理论严谨的新框架,挑战了当前以仿射+非线性为主导的范式,并建议将归一化重新理解为带参数缩放的激活函数类映射。
原文摘要 · Abstract (English)
A systematic mismatch exists between mathematically ideal and effective activation updates during gradient descent. As intended, parameters update in their direction of steepest descent. However, activations are argued to constitute a more directly impactful quantity to prioritise in optimisation, as they are closer to the loss in the computational graph and carry sample-dependent information through the network. Yet their propagated updates do not take the optimal steepest-descent step. These quantities exhibit non-ideal sample-wise scaling across affine, convolutional, and attention layers.Solutions to correct for this are trivial and, incidentally, derive normalisation from first principles despite motivational independence. Consequently, such considerations offer a fresh, conceptual reframe of normalisation's action, with auxiliary experiments bolstering this mechanistic interpretation. Moreover, this analysis makes clear a second possibility: a solution that is functionally distinct from modern normalisations, without scale invariance, yet remains empirically successful -- an alternative to the affine map. This outperforms conventional normalisers across several tests. This generalises to convolution via a new functional form, ``PatchNorm'', a compositionally inseparable normaliser. Together, these provide an alternative mechanistic framework that both adds to and counters some of the discussion of normalisation. Further, it is argued that normalisers are better decomposed into activation-function-like maps with parameterised scaling. Overall, this constitutes a theoretically principled approach that yields new functions with empirical validation and raises questions about the affine + nonlinear approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。