用图论设计结构化剪枝,让大模型更轻更快还保持精度。
EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models
- 基于膨胀图理论设计剪枝结构,保证信息流动
- 在多个大模型上实现显著加速与内存节省
- 适合追求高效部署的LLM研究人员和工程师
随着大语言模型规模不断增大,其部署带来的计算与内存挑战日益严峻,亟需更高效的模型变体。本文提出EGGS-PTP:一种基于膨胀图引导的结构化后训练剪枝方法。该方法利用图论指导N:M结构化剪枝设计,有效降低模型尺寸与计算开销。通过引入膨胀图概念,确保剪枝后网络中的信息流畅通,保留关键模型功能。大量实验表明,EGGS-PTP不仅因结构稀疏性实现显著加速与内存节省,还在多个大语言模型上超越现有结构化剪枝方法,在准确率方面表现更优。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) become more widely adopted and scale up in size, the computational and memory challenges involved in deploying these massive foundation models have grown increasingly severe. This underscores the urgent need to develop more efficient model variants. Faced with this challenge, the present work introduces EGGS-PTP: an Expander-Graph Guided Structured Post-training Pruning method. The proposed approach leverages graph theory to guide the design of N:M structured pruning, effectively reducing model size and computational demands. By incorporating concepts from expander graphs, EGGS-PTP ensures information flow within the pruned network, preserving essential model functionality. Extensive numerical experiments demonstrate that EGGS-PTP not only achieves significant acceleration and memory savings due to structured sparsity but also outperforms existing structured pruning techniques in terms of accuracy across various LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。