用高稀疏度与低精度量化,快速压缩大模型且保持高精度。
Aggressive Post-Training Compression on Extremely Large Language Models

- 采用超0.7稀疏度与低于8比特量化进行网络剪枝。
- 数小时内完成主流大模型压缩,精度损失小。
- 适合希望在本地设备部署大模型的开发者。
大型语言模型(LLMs)规模和复杂性的增加给其在个人电脑和移动设备上的部署带来了挑战。激进的训练后模型压缩对于减小模型体积至关重要,但通常会导致显著的精度下降。为应对这一挑战,我们提出一种新型网络剪枝技术,采用超过0.7的稀疏度和少于8比特的量化。该方法可在数小时内压缩当前主流的LLMs,同时保持相对较小的精度损失。实验评估表明,该方法具有有效性及实际部署潜力。通过使大模型可在本地设备上运行,本工作有望推动自然语言处理应用进入新纪元。
原文摘要 · Abstract (English)
The increasing size and complexity of Large Language Models (LLMs) pose challenges for their deployment on personal computers and mobile devices. Aggressive post-training model compression is necessary to reduce the models' size, but it often results in significant accuracy loss. To address this challenge, we propose a novel network pruning technology that utilizes over 0.7 sparsity and less than 8 bits of quantization. Our approach enables the compression of prevailing LLMs within a couple of hours while maintaining a relatively small accuracy loss. In experimental evaluations, our method demonstrates effectiveness and potential for practical deployment. By making LLMs available on domestic devices, our work can facilitate a new era of natural language processing applications with wide-ranging impacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。