arXiv:2605.07271cs.CLcs.AI2026-05被引 1

揭示大模型剪枝性能骤降的根源:关键在沉默阶段被破坏

Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions

论文配图:Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions
图 1 · 摘自论文原文
  • 通过决策表征分析剪枝过程,发现网络存在沉默与决断两个阶段
  • 剪掉决断阶段影响小,剪掉沉默阶段直接导致性能崩溃
  • 适合研究模型压缩与鲁棒性的人参考

层剪枝能有效降低大语言模型的计算开销,但常引发突然的性能下降。现有基于表示的分析难以解释其机制。本文从决策表征出发,聚焦多选任务,提出决策边界和选项频率两个度量,并设计迭代剪枝方法,分析逐层决策动态。研究发现,网络存在明显的决策跃迁,将模型分为两个阶段:沉默阶段(尚无法正确预测)和决断阶段(正确答案出现)。结果表明,剪掉决断阶段影响较小,而剪掉沉默阶段会立即引发性能崩溃,说明该阶段对结构变化极度敏感。因此,剪枝导致的性能下降源于对沉默阶段的破坏,阻止了关键决策跃迁的发生。

原文摘要 · Abstract (English)

Layer pruning efficiently reduces Large Language Model (LLM) computational costs but often triggers sudden performance collapse. Existing representation-based analyses struggle to explain this mechanism. We propose studying pruning through decision representation. Focusing on multiple-choice tasks, we introduce two metrics, Decision Margin and Option Frequency, and an Iterative Pruning method to analyze layer-wise decision dynamics. Our findings reveal a sharp decision transition that partitions the network into two stages: a Silent Phase, where the model cannot yet predict the correct answer, and a Decisive Phase, where the correct prediction emerges. We also find that pruning the Decisive Phase has minimal impact, whereas pruning the Silent Phase triggers immediate performance collapse, highlighting its extreme sensitivity to structural changes. Therefore, we conclude that pruning-induced collapse stems from disrupting the Silent Phase, which prevents the critical decision transition from occurring.

模型剪枝决策分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。