arXiv:2409.00492cs.CV2024-09被引 3

用向量量化压缩文生图模型,3比特仍保高质量

Accurate Compression of Text-to-Image Diffusion Models via Vector Quantization

  • 采用向量量化技术替代传统量化方法
  • 20亿参数模型压缩至约3比特,图像质量接近4比特方案
  • 适合资源受限环境下部署大模型

文生图扩散模型已成为基于文本提示生成高质量图像的强大框架。其成功推动了生产级扩散模型的快速发展,模型规模持续增长,现已包含数十亿参数。因此,当前最先进的文生图模型在实际应用中变得越来越难以访问,尤其是在资源受限环境中。后训练量化(PTQ)通过将预训练模型权重压缩为低比特表示来解决此问题。现有的扩散模型量化技术主要依赖于均匀标量量化,在压缩至4比特时表现良好。本文表明,更灵活的向量量化(VQ)可为大规模文生图扩散模型实现更高的压缩率。具体而言,我们将基于向量的PTQ方法适配到最新的数十亿参数文生图模型(SDXL和SDXL-Turbo),并证明使用VQ将20亿参数以上的模型压缩至约3比特时,仍能保持与此前4比特压缩技术相当的图像质量和文本对齐效果。

原文摘要 · Abstract (English)

Text-to-image diffusion models have emerged as a powerful framework for high-quality image generation given textual prompts. Their success has driven the rapid development of production-grade diffusion models that consistently increase in size and already contain billions of parameters. As a result, state-of-the-art text-to-image models are becoming less accessible in practice, especially in resource-limited environments. Post-training quantization (PTQ) tackles this issue by compressing the pretrained model weights into lower-bit representations. Recent diffusion quantization techniques primarily rely on uniform scalar quantization, providing decent performance for the models compressed to 4 bits. This work demonstrates that more versatile vector quantization (VQ) may achieve higher compression rates for large-scale text-to-image diffusion models. Specifically, we tailor vector-based PTQ methods to recent billion-scale text-to-image models (SDXL and SDXL-Turbo), and show that the diffusion models of 2B+ parameters compressed to around 3 bits using VQ exhibit the similar image quality and textual alignment as previous 4-bit compression techniques.

文生图模型压缩向量量化扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。