低比特量化让大模型更省资源,这篇综述系统梳理了方法与系统。
A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms
- 从基础原理到系统实现,全面梳理低比特量化技术
- 涵盖训练与推理优化工具,支持多硬件平台部署
- 适合想提升大模型效率的研究者和工程师参考
大语言模型(LLMs)在自然语言处理中取得显著进展,但在实际部署中面临高昂的内存与计算成本。低比特量化通过降低模型参数、激活值和梯度的位宽,有效减少内存占用与计算需求,成为关键解决方案。本文系统综述面向LLMs的低比特量化方法,涵盖基本原理、系统实现与算法策略。首先介绍低比特LLM的基础概念与新型数据格式,随后回顾跨硬件平台支持低比特LLM的框架与系统。接着对高效低比特训练与推理的技术及工具包进行分类分析。最后讨论未来趋势与潜在发展方向。从基础、系统、算法三方面提供的系统性综述,为通过低比特量化提升LLM效率与可用性提供重要参考。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved remarkable advancements in natural language processing, showcasing exceptional performance across various tasks. However, the expensive memory and computational requirements present significant challenges for their practical deployment. Low-bit quantization has emerged as a critical approach to mitigate these challenges by reducing the bit-width of model parameters, activations, and gradients, thus decreasing memory usage and computational demands. This paper presents a comprehensive survey of low-bit quantization methods tailored for LLMs, covering the fundamental principles, system implementations, and algorithmic strategies. An overview of basic concepts and new data formats specific to low-bit LLMs is first introduced, followed by a review of frameworks and systems that facilitate low-bit LLMs across various hardware platforms. Then, we categorize and analyze techniques and toolkits for efficient low-bit training and inference of LLMs. Finally, we conclude with a discussion of future trends and potential advancements of low-bit LLMs. Our systematic overview from basic, system, and algorithm perspectives can offer valuable insights and guidelines for future works to enhance the efficiency and applicability of LLMs through low-bit quantization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。