arXiv:2504.14701cs.LGstat.ML2025-04中稿 · TMLR 2025被引 2

发现大参数方向与损失曲率主空间高度重合,揭示了深度学习优化的几何本质。

Connecting Parameter Magnitudes and Hessian Eigenspaces at Scale using Sketched Methods

  • 用格拉斯曼距离度量参数掩码与海森矩阵特征空间的重合度
  • 在超1000万参数模型上实现上千个海森特征对的高效计算
  • 大参数更易对齐高曲率方向,适用于理解模型结构与压缩机制

最近研究发现,使用SGD训练深度神经网络时,损失曲面的大部分曲率迅速集中于海森矩阵的极小特征子空间,且后续保持稳定。同时,成功的参数剪枝掩码也早在训练早期形成并趋于稳定。本文首次联合研究这两类现象,提出基于格拉斯曼距离的相似性度量方法,发现‘重合度’最具可解释性与稳定性。为计算该指标,我们开发了一种基于压缩奇异值分解(sketched SVD)的无矩阵算法,可在超过1000万参数的网络上高效计算超过1000个海森特征对——这是前所未有的规模。实验表明,参数大小掩码与顶部海森特征空间的重合度显著高于随机水平,且随网络增大而增强。结果表明:海森矩阵的主特征向量倾向于集中在较大参数方向,即大参数更倾向于对齐高损失曲率方向。本工作提供了一种大规模近似分析深度学习海森矩阵的方法,并揭示其特征空间结构的新洞见。

原文摘要 · Abstract (English)

Recently, it has been observed that when training a deep neural net with SGD, the majority of the loss landscape's curvature quickly concentrates in a tiny *top* eigenspace of the loss Hessian, which remains largely stable thereafter. Independently, it has been shown that successful magnitude pruning masks for deep neural nets emerge early in training and remain stable thereafter. In this work, we study these two phenomena jointly and show that they are connected: We develop a methodology to measure the similarity between arbitrary parameter masks and Hessian eigenspaces via Grassmannian metrics. We identify *overlap* as the most useful such metric due to its interpretability and stability. To compute *overlap*, we develop a matrix-free algorithm based on sketched SVDs that allows us to compute over 1000 Hessian eigenpairs for nets with over 10M parameters --an unprecedented scale by several orders of magnitude. Our experiments reveal an *overlap* between magnitude parameter masks and top Hessian eigenspaces consistently higher than chance-level, and that this effect gets accentuated for larger network sizes. This result indicates that *top Hessian eigenvectors tend to be concentrated around larger parameters*, or equivalently, that *larger parameters tend to align with directions of larger loss curvature*. Our work provides a methodology to approximate and analyze deep learning Hessians at scale, as well as a novel insight on the structure of their eigenspace.

海森矩阵参数重要性特征分析大规模优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。