arXiv:2505.02809cs.LGmath.OC2025-05被引 13

揭示神经网络海森矩阵近块对角结构的根源,发现类别数是关键因素。

Towards Quantifying the Hessian Structure of Neural Networks

  • 从架构设计与训练动态双角度解析海森矩阵结构
  • 当分类类别数 $C$ 很大时,海森矩阵呈明显块对角分布
  • 适用于研究大语言模型等高类别场景的优化机制

实证研究表明,神经网络的海森矩阵呈现近块对角结构,但其理论基础尚不明确。本文揭示该结构源于两种力的共同作用:由网络架构决定的‘静态力’和由训练过程引发的‘动态力’。我们对随机初始化下的‘静态力’进行严格理论分析,研究线性模型及单隐层分类网络($C$ 类)。基于随机矩阵理论,比较对角与非对角海森块的极限分布,发现当 $C$ 较大时,块对角结构自然出现。结果表明,$C$ 是导致近块对角结构的主要驱动力。该发现或可为大规模语言模型(通常 $C > 10^4$)的海森结构提供新视角。

原文摘要 · Abstract (English)

Empirical studies reported that the Hessian matrix of neural networks (NNs) exhibits a near-block-diagonal structure, yet its theoretical foundation remains unclear. In this work, we reveal that the reported Hessian structure comes from a mixture of two forces: a ``static force'' rooted in the architecture design, and a ''dynamic force'' arisen from training. We then provide a rigorous theoretical analysis of ''static force'' at random initialization. We study linear models and 1-hidden-layer networks for classification tasks with $C$ classes. By leveraging random matrix theory, we compare the limit distributions of the diagonal and off-diagonal Hessian blocks and find that the block-diagonal structure arises as $C$ becomes large. Our findings reveal that $C$ is one primary driver of the near-block-diagonal structure. These results may shed new light on the Hessian structure of large language models (LLMs), which typically operate with a large $C$ exceeding $10^4$.

神经网络海森矩阵随机矩阵分类任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。