发现量化与稀疏性在模型缩放中具有统一规律,可高效压缩大模型。
Compression Scaling Laws:Unifying Sparsity and Quantization
- 提出统一的缩放定律框架,涵盖权重稀疏与量化。
- 仅量化权重时效率高,全量化的收益随位宽降低而递减。
- 适合研究模型压缩与高效训练的学者参考。
我们研究了不同压缩技术——如权重和激活量化、权重稀疏性——对大语言模型预训练过程中缩放行为的影响。基于已有工作表明权重稀疏性在缩放定律中表现为模型规模的恒定倍数,我们进一步证明该‘有效参数’缩放模式同样适用于量化。具体而言,仅权重量化能实现强参数效率倍数,而权重与激活全量化在低比特位宽下表现出收益递减。结果表明,不同压缩技术可统一于同一缩放定律框架,支持方法间的合理比较与组合。
原文摘要 · Abstract (English)
We investigate how different compression techniques -- such as weight and activation quantization, and weight sparsity -- affect the scaling behavior of large language models (LLMs) during pretraining. Building on previous work showing that weight sparsity acts as a constant multiplier on model size in scaling laws, we demonstrate that this "effective parameter" scaling pattern extends to quantization as well. Specifically, we establish that weight-only quantization achieves strong parameter efficiency multipliers, while full quantization of both weights and activations shows diminishing returns at lower bitwidths. Our results suggest that different compression techniques can be unified under a common scaling law framework, enabling principled comparison and combination of these methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。