揭示批量归一化如何延迟损失突增的内在机制
A Mechanism Study of Delayed Loss Spikes in Batch-Normalized Linear Models

- 通过分析归一化线性模型,发现归一化会逐步提升有效学习率
- 在白化平方损失下,证明了延迟突增的起始条件与稳定边界
- 为理解训练不稳定性提供具体路径,适合研究优化机制者
训练中出现的延迟损失突增现象已有报道,但现有理论主要解释由过大固定学习率引发的早期非单调行为。本文提出一个简化假设:归一化可通过在原本稳定的下降过程中逐渐增加有效学习率,从而推迟不稳定性。为在定理层面验证该假设,研究了批量归一化的线性模型。核心结果针对白化平方损失线性回归,推导出无上升边缘和延迟起始的条件,界定了方向性突增的等待时间,并证明上升边缘可在有限步内自我稳定。结合平方损失分解,明确揭示了白化情形下的延迟突增机制。对于逻辑回归,在高度受限的活跃间隔假设下,仅在临界区域证明了有限时域的方向性前兆,附录中额外给出在非退化条件下的损失下界。因此,本文应被视为一种简化机制研究,而非对神经网络损失突增的普遍解释。在此范围内,结果明确识别出由批量归一化引发的一种具体延迟不稳定性路径。
原文摘要 · Abstract (English)
Delayed loss spikes have been reported in neural-network training, but existing theory mainly explains earlier non-monotone behavior caused by overly large fixed learning rates. We study one stylized hypothesis: normalization can postpone instability by gradually increasing the effective learning rate during otherwise stable descent. To test this hypothesis at theorem level, we analyze batch-normalized linear models. Our flagship result concerns whitened square-loss linear regression, where we derive explicit no-rising-edge and delayed-onset conditions, bound the waiting time to directional onset, and show that the rising edge self-stabilizes within finitely many iterations. Combined with a square-loss decomposition, this yields a concrete delayed-spike mechanism in the whitened regime. For logistic regression, under highly restrictive active-margin assumptions, we prove only a supporting finite-horizon directional precursor in a knife-edge regime, with an optional appendix-only loss lower bound under an extra non-degeneracy condition. The paper should therefore be read as a stylized mechanism study rather than a general explanation of neural-network loss spikes. Within that scope, the results isolate one concrete delayed-instability pathway induced by batch normalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。