通过规范固定思想改善神经网络训练稳定性,提升学习率范围。
Scale redundancy and soft gauge fixing in positively homogeneous neural networks
- 引入软轨道选择项,自动平衡神经元尺度以消除冗余自由度。
- 实验表明可扩大稳定学习率范围,抑制尺度漂移且不损失模型表达力。
- 为机器学习优化提供了规范场论的理论新视角,适合研究优化与结构的学者。
具有正齐次激活函数的神经网络存在精确的连续重参数化对称性:单个神经元的缩放不会改变输入-输出函数。我们将其解释为规范冗余,并引入规范适应坐标,分离不变量与尺度不平衡方向。受场论中规范固定启发,提出仅作用于冗余尺度坐标的软轨道选择(范数平衡)正则项。理论上证明该正则项能诱导不平衡模式的耗散性松弛,从而保持实现函数不变。在受控实验中,该方法扩展了稳定学习率区间,抑制了尺度漂移,且未改变模型表达能力。结果建立了规范轨道几何与优化条件之间的结构性联系,为规范场论概念与机器学习提供了具体关联。
原文摘要 · Abstract (English)
Neural networks with positively homogeneous activations exhibit an exact continuous reparametrization symmetry: neuron-wise rescalings generate parameter-space orbits along which the input--output function is invariant. We interpret this symmetry as a gauge redundancy and introduce gauge-adapted coordinates that separate invariant and scale-imbalance directions. Inspired by gauge fixing in field theory, we introduce a soft orbit-selection (norm-balancing) functional acting only on redundant scale coordinates. We show analytically that it induces dissipative relaxation of imbalance modes to preserve the realized function. In controlled experiments, this orbit-selection penalty expands the stable learning-rate regime and suppresses scale drift without changing expressivity. These results establish a structural link between gauge-orbit geometry and optimization conditioning, providing a concrete connection between gauge-theoretic concepts and machine learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。