首个系统性评测小模型量化,揭示其与大模型的压缩差异。
SLMQuant:Benchmarking Small Language Model Quantization for Practical Deployment
- 构建 SLMQuant 基准,跨架构多任务评估小模型量化效果。
- 发现小模型对量化更敏感,直接套用大模型方法效果差。
- 提出适配小模型的压缩设计原则,适合边缘部署场景。
尽管小型语言模型(SLMs)作为资源高效替代方案备受关注,但其在边缘设备上的部署仍面临模型压缩效率不足的挑战。虽然量化在大型语言模型(LLMs)中已证明有效,但其在小模型中的适用性却显著未被探索,关键问题包括量化瓶颈差异和效率特征不同。本文提出 SLMQuant,首个系统性评估将大模型压缩技术应用于小模型的基准。通过跨多种架构和任务的多维度评估,分析先进量化方法在小模型上的表现。研究发现,小模型与大模型在量化敏感性上存在根本差异,直接迁移针对大模型优化的技术会导致次优结果,原因在于小模型独特的架构特征和训练动态。我们识别出影响小模型有效量化的关键因素,并提出可操作的设计原则以实现小模型定制化压缩。SLMQuant 为推进低功耗设备上的高效小模型部署奠定了基础,并为资源受限场景下的轻量级语言模型部署提供了关键洞见。
原文摘要 · Abstract (English)
Despite the growing interest in Small Language Models (SLMs) as resource-efficient alternatives to Large Language Models (LLMs), their deployment on edge devices remains challenging due to unresolved efficiency gaps in model compression. While quantization has proven effective for LLMs, its applicability to SLMs is significantly underexplored, with critical questions about differing quantization bottlenecks and efficiency profiles. This paper introduces SLMQuant, the first systematic benchmark for evaluating LLM compression techniques when applied to SLMs. Through comprehensive multi-track evaluations across diverse architectures and tasks, we analyze how state-of-the-art quantization methods perform on SLMs. Our findings reveal fundamental disparities between SLMs and LLMs in quantization sensitivity, demonstrating that direct transfer of LLM-optimized techniques leads to suboptimal results due to SLMs' unique architectural characteristics and training dynamics. We identify key factors governing effective SLM quantization and propose actionable design principles for SLM-tailored compression. SLMQuant establishes a foundational framework for advancing efficient SLM deployment on low-end devices in edge applications, and provides critical insights for deploying lightweight language models in resource-constrained scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。