arXiv:2609.06557cs.LG2026-09

通过重访基础剪枝策略,实现大模型99%稀疏下的高性能。

Hidden in Plain Sight: The Overlooked Significance of Canonical Elements for Extreme LLM Sparsity

论文配图:Hidden in Plain Sight: The Overlooked Significance of Canonical Elements for Extreme LLM Sparsity
图 1 · 摘自论文原文
  • 采用二阶重要性与渐进训练配合的剪枝框架。
  • 在95%和99%稀疏下,困惑度分别降至13.48和19.67。
  • 适合追求极致压缩与推理加速的模型部署场景。

大语言模型(LLMs)常被认为在激进稀疏化下易失稳,维持可靠性能通常需限制在中等稀疏水平。然而,近期研究显示LLMs对高稀疏更具韧性,问题应被重新视为设计挑战而非根本限制。本文通过重构未充分探索的基础剪枝策略,提出一种结合二阶重要性与持续训练的渐进稀疏化框架。在LLaMA-2与Qwen-3系列模型上,该方法可将预训练模型性能保持在极端稀疏区间。具体而言,在LLaMA-2-7B上,95%和99%稀疏下的WikiText-2困惑度分别为13.48和19.67,同时实现3.23倍解码加速与6.21倍内存节省。结果表明,大型语言模型可在极强稀疏下仍保持良好性能,为该稀疏区间的模型优化奠定基础。

原文摘要 · Abstract (English)

Large language models (LLMs) are often considered fragile under aggressive sparsification, and maintaining reliable performance typically requires sticking to moderate sparsity levels. However, recent studies suggest that LLMs are more resilient to high sparsity than previously thought, reframing the problem as a design challenge rather than a fundamental limitation. In this work, we challenge the perceived limits of unstructured post-training LLM pruning by revisiting elementary pruning strategies that have remained relatively underexplored at this scale. Through a progressive sparsification framework with second-order saliency and continued training coordinated with sparsity progression, we show that pretrained LLMs can retain strong performance far beyond commonly studied sparsity regimes. Across LLaMA-2 and Qwen-3 model families, our approach improves perplexity and downstream accuracy up to 99\% sparsity, surpassing both the current state-of-the-art and representative baselines. Precisely, on LLaMA-2-7B, our approach achieves WikiText-2 perplexities of 13.48 and 19.67 at 95\% and 99\% sparsity, respectively, while delivering 3.23$\times$ decoding speedup and 6.21$\times$ memory savings at 95\% sparsity. Taken together, our results show that LLMs can be pushed into extreme sparsity while retaining strong performance, providing a foundation for further improving sparse models in this regime.

模型压缩稀疏化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。