提出新方法精准测量神经网络训练中的几何结构变化。
Bypassing Minimization Bias: A Shift-Invariant Variance Estimator for Off-Equilibrium Local Learning Coefficients
- 用方差操作消除基准值影响,避免偏差。
- 在非平衡训练阶段仍能准确捕捉几何信号。
- 适合研究模型结构演化与训练动态的学者。
奇异学习理论利用局部学习系数(LLC)量化神经网络损失曲面的几何特性。然而,基于均值能量的LLC估计依赖于一个加性损失基准,通常为局部极小值的估计。在瞬态、非平衡的训练阶段,该极小值未知;若以最小噪声小批量损失替代,会引入系统性最小化偏差,扭曲几何度量。本文提出移位不变方差估计器(SIVE),一种基于方差的局部LLC探测方法,通过方差算子结构上消除未知的加性基准。结合从全方差定律导出的显式修正,SIVE将几何损失波动与小批量评估噪声分离。在可解析处理的模型上进行受控实验表明,当锚定均值估计器失效时,SIVE仍能恢复预期的有限温度几何信号。应用于深度神经网络时,SIVE提供了一种鲁棒的、局部化的在线诊断工具,可用于全程追踪训练过程中的结构相变。
原文摘要 · Abstract (English)
Singular Learning Theory leverages the Local Learning Coefficient (LLC) to quantify the geometry of neural network loss landscapes. However, mean-energy LLC estimators depend explicitly on an additive loss baseline, typically an estimate of the local minimum. During transient, off-equilibrium training phases, this minimum is unknown; substituting it with the lowest noisy mini-batch loss induces a systematic minimization bias that distorts the geometric measurement. In this paper, we propose the Shift-Invariant Variance Estimator (SIVE), a variance-based local LLC probe that structurally eliminates the unknown additive baseline through the variance operator. Combining this shift-invariant observable with an explicit correction derived from the Law of Total Variance, SIVE separates geometric loss fluctuations from mini-batch evaluation noise. Controlled experiments on analytically tractable toy models show that SIVE recovers the expected finite-temperature geometric signal in regimes where anchored mean estimators fail. Applied to deep neural networks, SIVE provides a robust, localized online diagnostic for tracking structural phase transitions throughout training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。