基于率失真理论的LLM压缩方法,可灵活控制模型大小与精度。
Radio: Rate-Distortion Optimization for Large Language Model Compression
- 从率失真理论出发设计量化方法,兼顾压缩效率与性能。
- 支持百亿参数模型的后训练压缩,用户可自定义模型大小与准确率。
- 适合需要在低资源设备部署大模型的研究者与开发者。
近年来,大语言模型(LLMs)的压缩已成为推动其在资源受限设备上部署、降低计算成本及缓解大规模人工智能基础设施环境影响的关键问题。本文从率失真理论的角度建立LLM量化的基础,并提出一种基于简单率失真优化的量化技术。该技术可扩展至包含数百亿参数的模型,支持用户在后训练阶段灵活压缩模型至指定的模型尺寸或精度要求。
原文摘要 · Abstract (English)
In recent years, the compression of large language models (LLMs) has emerged as a key problem in facilitating LLM deployment on resource-limited devices, reducing compute costs, and mitigating the environmental footprint due to large-scale AI infrastructure. Here, we establish the foundations of LLM quantization from a rate-distortion theory perspective and propose a quantization technique based on simple rate-distortion optimization. Our technique scales to models containing hundreds of billions of weight parameters and offers users the flexibility to compress models, post-training, to a model size or accuracy specified by the user.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。