提出可量化敏感度分析框架,优化扩散模型混合精度压缩效果
Qua$^2$SeDiMo: Quantifiable Quantization Sensitivity of Diffusion Models
- 构建混合精度量化框架,分析不同层与结构的量化敏感度
- 实现PixArt-α等模型3.4~3.7比特权重量化,优于现有方法
- 适合需要高效部署扩散模型的研究者与工程师
扩散模型(DM)通过迭代去噪过程实现了大众化的图像生成。量化是降低推理成本、减小去噪器网络尺寸的关键技术。然而,随着去噪器从卷积U-Net架构演变为新型Transformer架构,理解不同权重层、操作和架构类型对量化敏感度的影响变得愈发重要。本文提出Qua$^2$SeDiMo,一个可解释的混合精度后训练量化框架,能针对不同去噪器操作类型和模块结构提供量化成本效益分析。基于这些洞察,我们为从基础U-Net到最先进Transformer的多种扩散模型制定高质量的混合精度量化方案。实验结果表明,该方法在PixArt-α、PixArt-Σ、Hunyuan-DiT和SDXL上分别实现了3.4比特、3.9比特、3.65比特和3.7比特的权重量化,并与6比特激活量化结合,在定量指标和生成图像质量上均超越现有方法。
原文摘要 · Abstract (English)
Diffusion Models (DM) have democratized AI image generation through an iterative denoising process. Quantization is a major technique to alleviate the inference cost and reduce the size of DM denoiser networks. However, as denoisers evolve from variants of convolutional U-Nets toward newer Transformer architectures, it is of growing importance to understand the quantization sensitivity of different weight layers, operations and architecture types to performance. In this work, we address this challenge with Qua$^2$SeDiMo, a mixed-precision Post-Training Quantization framework that generates explainable insights on the cost-effectiveness of various model weight quantization methods for different denoiser operation types and block structures. We leverage these insights to make high-quality mixed-precision quantization decisions for a myriad of diffusion models ranging from foundational U-Nets to state-of-the-art Transformers. As a result, Qua$^2$SeDiMo can construct 3.4-bit, 3.9-bit, 3.65-bit and 3.7-bit weight quantization on PixArt-$α$, PixArt-$Σ$, Hunyuan-DiT and SDXL, respectively. We further pair our weight-quantization configurations with 6-bit activation quantization and outperform existing approaches in terms of quantitative metrics and generative image quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。