arXiv:2505.02214cs.LG2025-05被引 28

测试Qwen3在1-8比特量化下的表现,发现低精度时语言任务明显退化

An Empirical Study of Qwen3 Quantization

  • 系统评估5种后训练量化方法在1-8比特范围的表现
  • 4比特以上仍保持较好性能,1-2比特时语言任务显著下降
  • 为Qwen3高效部署提供实证参考,适合模型压缩研究者

Qwen系列已成为领先的开源大语言模型家族,在自然语言理解任务中表现出色。随着Qwen3的发布,其在多个基准测试中展现出卓越性能,引发对在资源受限环境中高效部署的关注。低比特量化是可行方案,但其对Qwen3的影响尚未充分研究。本研究系统评估了5种经典后训练量化技术在Qwen3上的表现,覆盖1至8比特多种位宽,并在多个数据集上进行测试。结果表明,尽管在中等位宽下仍具竞争力,但在超低精度(如1-2比特)下,语言任务性能显著下降,凸显大模型压缩中的持续挑战。该分析为优化针对Qwen3及未来LLM的量化方法提供了实践指导,有助于提升模型实用性而不牺牲准确性。项目代码与模型已公开于GitHub和HuggingFace。

原文摘要 · Abstract (English)

The Qwen series has emerged as a leading family of open-source Large Language Models (LLMs), demonstrating remarkable capabilities in natural language understanding tasks. With the recent release of Qwen3, which exhibits superior performance across diverse benchmarks, there is growing interest in deploying these models efficiently in resource-constrained environments. Low-bit quantization presents a promising solution, yet its impact on Qwen3's performance remains underexplored. This study conducts a systematic evaluation of Qwen3's robustness under various quantization settings, aiming to uncover both opportunities and challenges in compressing this state-of-the-art model. We rigorously assess 5 existing classic post-training quantization techniques applied to Qwen3, spanning bit-widths from 1 to 8 bits, and evaluate their effectiveness across multiple datasets. Our findings reveal that while Qwen3 maintains competitive performance at moderate bit-widths, it experiences notable degradation in linguistic tasks under ultra-low precision, underscoring the persistent hurdles in LLM compression. These results emphasize the need for further research to mitigate performance loss in extreme quantization scenarios. We anticipate that this empirical analysis will provide actionable insights for advancing quantization methods tailored to Qwen3 and future LLMs, ultimately enhancing their practicality without compromising accuracy. Our project is released on https://github.com/Efficient-ML/Qwen3-Quantization and https://huggingface.co/collections/Efficient-ML/qwen3-quantization-68164450decb1c868788cb2b.

模型压缩量化Qwen3

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。