arXiv:2602.00816stat.MLcs.LG2026-02被引 1

首次实现百亿参数大模型的精准海森谱分析,揭示现有近似方法严重失真。

Hessian Spectral Analysis at Foundation Model Scale

  • 采用分片局部有限差分法,兼容全分片数据并行,实现大规模海森向量乘积。
  • 在100B参数语言模型上首次获得超100亿规模的海森谱密度估计结果。
  • 发现块对角曲率近似误差可达一阶,适用于需精确曲率分析的研究者。

基础模型的真实海森谱分析长期难以实现,以往研究多依赖小模型或强结构假设。本文展示在前沿规模下进行真实海森谱分析是可行的。通过与全分片数据并行兼容的分片本地有限差分海森向量乘积,我们在开源语言模型中实现了高达1000亿参数的随机兰佐斯求积,首次获得超越100亿参数规模的大型谱密度估计。我们刻画了该流水线的数值行为,包括有限差分偏差、浮点噪声放大及其对fp32和bf16中克里洛夫稳定性的影响,并推导出经实证验证的实用操作区间。进一步提供了端到端的运行时与内存缩放规律,表明全算子谱探测仅带来训练的一阶常数级开销。关键的是,直接访问海森矩阵揭示,广泛使用的块对角曲率近似会灾难性失效,即使在中等规模大模型中也表现出一阶相对误差和较差的方向一致性。综合来看,我们的结果证明基础模型的海森谱既可计算,又普遍被现有近似方法严重误判,为大规模曲率驱动分析开辟了新路径。

原文摘要 · Abstract (English)

Accurate Hessian spectra of foundation models have remained out of reach, leading most prior work to rely on small models or strong structural approximations. We show that faithful spectral analysis of the true Hessian is tractable at frontier scale. Using shard-local finite-difference Hessian vector products compatible with Fully Sharded Data Parallelism, we perform stochastic Lanczos quadrature on open-source language models with up to 100B parameters, producing the first large-scale spectral density estimates beyond the sub-10B regime. We characterize the numerical behavior of this pipeline, including finite-difference bias, floating-point noise amplification, and their effect on Krylov stability in fp32 and bf16, and derive practical operating regimes that are validated empirically. We further provide end-to-end runtime and memory scaling laws, showing that full-operator spectral probing incurs only a modest constant-factor overhead over first-order training. Crucially, direct access to the Hessian reveals that widely used block-diagonal curvature approximations can fail catastrophically, exhibiting order-one relative error and poor directional alignment even in mid-scale LLMs. Together, our results demonstrate that foundation-model Hessian spectra are both computable and qualitatively misrepresented by prevailing approximations, opening the door to principled curvature-based analysis at scale.

海森谱大模型曲率分析语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。