arXiv:2606.15161cs.CL2026-06被引 1

提出层间扰动吸收机制,提升大模型稀疏化压缩效果

Beyond Layer Importance in Layer-wise Sparsity: An Inter-Layer Perturbation-Absorption Perspective

论文配图:Beyond Layer Importance in Layer-wise Sparsity: An Inter-Layer Perturbation-Absorption Perspective
图 1 · 摘自论文原文
  • 从层间扰动吸收角度重新理解稀疏化,发现中后层能主动吸收误差
  • 在70%稀疏度下,困惑度降低7.13%,零样本准确率提升1.02%
  • 适合关注模型压缩效率与鲁棒性的研究者

大语言模型存在显著的层间冗余,非均匀稀疏分配已成为高效压缩的标准方法。现有基于局部信号(如激活异常或权重谱)估计层重要性的方法主要依赖局部层重要性,但最终压缩性能还受网络后续补偿能力影响。本文通过受控扰动实验直接刻画这一特性。实证发现:层对剪枝规模扰动响应高度异质;多数情况下,早期层放大扰动,中后层主动吸收扰动,相对L2漂移随深度单调下降,方向逐渐回归未扰动隐藏状态轨迹。吸收是大扰动现象:小扰动时全层放大,随扰动增大至剪枝尺度,吸收转变平滑发生。基于此,我们定义每层吸收系数,提出吸收感知修正方法,在多个模型族70%稀疏度下,使OWL和AlphaPruning的困惑度降低7.13%,零样本准确率提升1.02%。

原文摘要 · Abstract (English)

The considerable layer-wise redundancy in large language models (LLMs) has established non-uniform sparsity allocation across layers as the standard pruning approach for efficient compression. Existing layer-wise allocation methods that estimate allocation strategy from local signals such as activation outliers or weight spectra mainly derive from local layer importance, whereas the final post-pruning performance is also influenced by the network's subsequent compensatory capacity. In this paper, we directly characterize this property through controlled perturbation experiments. We make the following empirical findings. First, layers exhibit highly heterogeneous responses to pruning-scale perturbations. In most cases, early layers amplify perturbations, while middle and late layers actively absorb them, with relative L2 drift decreasing monotonically across depth and direction realigning toward the unperturbed hidden-state trajectory. Second, absorption is a large-perturbation phenomenon. Under small perturbations the network exhibits amplification across all layers, and the transition to absorption occurs smoothly as perturbation magnitude grows to pruning scale. This enriches the linearized accumulation theory underlying related works. Building on these findings, we define an absorption coefficient per layer and propose absorption-aware correction, an orthogonal augmentation that improves OWL and AlphaPruning by reducing perplexity by 7.13% and boosting zero-shot accuracy by 1.02% across multiple model families at 70% sparsity.

模型压缩稀疏化层间机制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。