提出高效近似方法,分析百亿参数模型的敏感层与脆弱结构。
Scalable Kronecker-Fisher Approximation: Efficient Hessian Analysis for Billion-Parameter Language Models Compression

- 用克罗内克分解近似海森矩阵,无需存储完整费雪信息。
- 发现价值投影层最敏感且跨层相关性强,影响压缩效果。
- 适用于量化、稀疏化等压缩策略,指导精度优化。
本文提出一种可扩展的基于克罗内克的近似方法,无需存储完整费雪矩阵即可捕捉跨层交互,使百亿参数网络的海森分析成为可能。实验表明,该方法揭示出一致的脆弱性模式:价值投影层对扰动最敏感且跨层相关性最强,其他组件则表现出架构特异性行为。在量化、稀疏化、层间扰动及扰动后微调等任务中,我们的近似结果与性能下降和恢复程度高度相关。该框架为大型模型中脆弱模块的识别提供了实用且理论严谨的工具,推动了混合精度分配、分层稀疏性及跨层甚至单权重组自适应低秩分解等定向压缩与优化策略的发展。
原文摘要 · Abstract (English)
In this paper, we propose a scalable Kronecker-based approximation that captures cross-layer interactions without storing the entire Fisher matrix, enabling practical Hessian analysis for billion-parameter networks where full computation is infeasible. Our approach reveals consistent vulnerability patterns: value projection layers exhibit the highest sensitivity and strongest cross-layer correlations across multiple model families, while other components exhibit architecture-specific behaviors. Through extensive experiments on quantization, sparsification, inter-layer corruption, and post-corruption fine-tuning, we demonstrate that our approximation strongly correlates with both performance degradation and recovery. Our framework provides a practical, theoretically grounded tool for identifying fragile components in large models, opening new avenues for guided compression and optimization strategies, such as mixed-precision allocation, layer-wise sparsity, and adaptive low-rank decomposition across layers and even individual weight groups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。