用费舍尔几何重新定义平坦性,解释为何平缓的极小值泛化更好。
Fisher-Geometric Sharpness and the Implicit Bias of SGD toward Flat Minima

- 基于费舍尔信息矩阵构建不变于参数重置的黎曼平坦度量
- 证明梯度噪声使模型收敛到黎曼平坦极小值,概率呈指数集中
- 实验验证该度量在MNIST和CIFAR-10上优于传统欧氏平坦度,且与学习率/批量大小匹配
深度学习中广泛认为随机梯度下降(SGD)隐式偏好平坦极小值,而平坦极小值具有更好的泛化能力。然而,标准的欧几里得平坦度度量(如损失海森矩阵的迹或最大特征值)在保持网络函数的前提下不满足参数重置下的不变性,削弱了这一观点的理论基础。本文通过将平坦性建立在由费舍尔信息矩阵(FIM)诱导的统计流形的黎曼几何之上,解决了该问题。我们数学定义了黎曼尖锐度,并证明其在光滑、函数保持的参数重置下不变,直接回应了Dinh等人在《尖锐极小值对深层网络也能泛化》中的质疑。值得注意的是,这一不变性仅严格存在于真实FIM;实践中使用的对角经验估计器仅近似具备该性质,而精确不变性需结构化估计器如K-FAC。我们形式化了小批量SGD的梯度噪声协方差结构与FIM成正比,推导出对应随机微分方程的平稳分布,并证明概率质量呈指数集中在黎曼平坦极小值处。一个由黎曼尖锐度(SR)显式控制的PAC-Bayes泛化界,将这种几何偏差与测试性能直接关联。在MNIST和CIFAR-10上的实验表明,SR能可靠追踪泛化性能,而欧氏平坦度则不能,且其在η/B下的缩放行为符合理论预测。这些结果共同提供了关于平坦极小值为何泛化的严格、参数不变的解释。
原文摘要 · Abstract (English)
A widely held intuition in deep learning is that stochastic gradient descent (SGD) implicitly favors flat minima and that flat minima generalize better, but standard Euclidean measures of flatness such as the trace or maximum eigenvalue of the loss Hessian are not invariant under reparametrizations that preserve the network function, which undermines the theoretical foundations of this narrative. In this study we resolve this issue by grounding flatness in the Riemannian geometry of the statistical manifold induced by the Fisher Information Matrix (FIM). We define Riemannian sharpness mathematically and prove that it is invariant under smooth, function-preserving reparametrizations, which directly addresses the critique of Dinh et al. in the paper ``Sharp minima can generalize for deep nets''.We note that this invariance is a property of the true FIM; the diagonal empirical estimator used in practice (and in all experiments below) inherits invariance only approximately, and exact invariance under arbitrary reparametrizations would require structured estimators such as K-FAC. We formalize the gradient noise of mini-batch SGD as having a covariance structure proportional to the FIM, derive the stationary distribution of the resulting stochastic differential equation, and then show that the probability mass is exponentially concentrated at Riemannian-flat minima. A PAC-Bayes generalization bound controlled explicitly by SR formally links this geometric bias to test performance. Our experiments on MNIST and CIFAR-10 confirm that SR reliably tracks generalization in ways that Euclidean sharpness does not, and that its scaling with $η/B$ matches the theoretical predictions. Together these results provide a rigorous, reparametrization-invariant account of why flat minima generalize.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。