arXiv:2506.12038cs.LGcs.AI2025-06被引 1

用知识蒸馏实现2-3比特超低精度聚类量化,大幅加速大模型推理。

LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation

  • 将聚类量化嵌入知识蒸馏框架,统一优化参数与激活
  • 2-3比特下保持大模型性能,推理速度最高提升6.2倍
  • 适合部署资源受限场景,如移动端或边缘设备

大型语言模型(LLMs)在自然语言处理中取得显著进展,但高内存和计算需求制约其部署。权重量化是缓解该问题的常用方法,但实现高效低比特压缩仍具挑战。本文提出LCD,将基于聚类的量化学习统一于知识蒸馏框架中。通过精心设计的优化技术,LCD在2-3比特超低精度下仍能保持LLM性能。此外,通过平滑策略压缩激活值,并采用查表(LUT)结构加速推理。实验表明,LCD优于现有方法,推理速度最高提升6.2倍。值得注意的是,LCD更具成本效益,适用于真实应用场景。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved significant progress in natural language processing but face challenges in deployment due to high memory and computational requirements. Weight quantization is a common approach to address these issues, yet achieving effective low-bit compression remains challenging. This paper presents LCD, which unifies the learning of clustering-based quantization within a knowledge distillation framework. Using carefully designed optimization techniques, LCD preserves LLM performance even at ultra-low bit widths of 2-3 bits. Additionally, LCD compresses activations through smoothing and accelerates inference with a LUT-based design. Experimental results show that LCD outperforms existing methods and delivers up to a 6.2x speedup in inference. Notably, LCD is shown to be more cost-effective, making it a practical solution for real-world applications.

大模型压缩低比特量化知识蒸馏推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。