提出对称剪枝新理论,显著提升大模型剪枝效果
Symmetric Pruning of Large Language Models
- 从输入激活与权重重要性双角度设计剪枝策略
- 新方法在多个数据集上超越现有基线,性能更优
- 无需训练的R²-DSnoT框架适合高效部署场景
主流的训练后剪枝方法如Wanda和RIA因其简单而高效的设计展现出卓越的实证性能。Wanda通过剪枝过程中的校准激活优化性能,RIA则强调权重元素的相对而非绝对重要性。尽管实践成功,但其理论基础仍不充分。本文提出新的理论洞察,重新定义剪枝的标准最小化目标,深入理解其成功因素。研究进一步提出兼顾输入激活与权重重要性的互补策略,并通过严格实验验证,显著优于现有方法。此外,我们引入一种新颖的无训练微调方法R²-DSnoT,结合相对权重重要性与正则化决策边界,在动态剪枝-生长框架中表现优异,大幅超越强基线,建立新基准。
原文摘要 · Abstract (English)
Popular post-training pruning methods such as Wanda and RIA are known for their simple, yet effective, designs that have shown exceptional empirical performance. Wanda optimizes performance through calibrated activations during pruning, while RIA emphasizes the relative, rather than absolute, importance of weight elements. Despite their practical success, a thorough theoretical foundation explaining these outcomes has been lacking. This paper introduces new theoretical insights that redefine the standard minimization objective for pruning, offering a deeper understanding of the factors contributing to their success. Our study extends beyond these insights by proposing complementary strategies that consider both input activations and weight significance. We validate these approaches through rigorous experiments, demonstrating substantial enhancements over existing methods. Furthermore, we introduce a novel training-free fine-tuning approach $R^2$-DSnoT that incorporates relative weight importance and a regularized decision boundary within a dynamic pruning-and-growing framework, significantly outperforming strong baselines and establishing a new state of the art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。