系统评估大模型量化在性能、能耗与质量上的权衡,指导实际部署。
Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
- 构建自动化工具qMeter,全面测试11种量化方法
- 发现量化效果高度依赖任务类型与硬件配置
- 适合关注大模型部署优化的研究者与工程师
大语言模型(LLMs)在多个领域表现出色,但其资源消耗巨大,量化(降低精度至低比特格式)成为高效服务的关键。尽管已有多种量化方法,但在真实服务场景下,其在性能、能耗与质量之间的系统性权衡仍不清晰。本文首先开发了全自动在线评估框架qMeter,对11种后训练量化方法在4种模型规模(7B-70B)和两种GPU架构(A100、H100)下进行了深入分析。评估覆盖应用层、工作负载、并行化和硬件层面的在线服务条件。研究揭示:量化效果高度依赖任务类型与方法,对工作负载特征敏感,并与并行策略和硬件架构存在复杂交互。进一步通过三个优化案例,展示了容量规划、能效调度与多目标调优中的部署挑战。据我们所知,这是首个从性能、能耗与质量联合视角,实现应用、系统与硬件层级全面表征的LLM量化研究。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable capabilities across diverse domains, but their heavy resource demands make quantization-reducing precision to lower-bit formats-critical for efficient serving. While many quantization methods exist, a systematic understanding of their performance, energy, and quality tradeoffs in realistic serving conditions remains a gap. In this work, we first develop a fully automated online characterization framework qMeter, and then conduct an in-depth characterization of 11 post-training LLM quantization methods across 4 model sizes (7B-70B) and two GPU architectures (A100, H100). We evaluate quantization at the application, workload, parallelism, and hardware levels under online serving conditions. Our study reveals highly task- and method-dependent tradeoffs, strong sensitivity to workload characteristics, and complex interactions with parallelism and GPU architecture. We further present three optimization case studies illustrating deployment challenges in capacity planning, energy-efficient scheduling, and multi-objective tuning. To the best of our knowledge, this is one of the first comprehensive application-, system-, and hardware-level characterization of LLM quantization from a joint performance, energy, and quality perspective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。