arXiv:2503.06564cs.CV2025-03被引 22

提出TR-DQ量化方法,提升扩散模型生成速度与内存效率

TR-DQ: Time-Rotation Diffusion Quantization

  • 按时间步分段,用旋转矩阵动态平滑激活和权重
  • 引入自适应超参数,实现不同时间步的动态量化
  • 在图像与视频生成上提速1.38-1.89倍,内存减少1.97-2.58倍

扩散模型在图像与视频生成中广泛应用,但其复杂结构导致推理开销高。现有量化方法主要关注模型结构量化,忽略采样过程中时间步变化的影响,且未能处理无法消除的重要激活,导致量化后性能显著下降。为此,我们提出时间-旋转扩散量化(TR-DQ),通过时间步划分并引入旋转矩阵,动态平滑激活与权重。针对不同时间步设计专用超参数,实现跨时间步的自适应量化。同时探索分类器自由引导(CFG-wise)的压缩潜力,为后续研究奠定基础。TR-DQ在图像与视频生成任务上达到当前最优性能,推理速度提升1.38-1.89倍,内存降低1.97-2.58倍。

原文摘要 · Abstract (English)

Diffusion models have been widely adopted in image and video generation. However, their complex network architecture leads to high inference overhead for its generation process. Existing diffusion quantization methods primarily focus on the quantization of the model structure while ignoring the impact of time-steps variation during sampling. At the same time, most current approaches fail to account for significant activations that cannot be eliminated, resulting in substantial performance degradation after quantization. To address these issues, we propose Time-Rotation Diffusion Quantization (TR-DQ), a novel quantization method incorporating time-step and rotation-based optimization. TR-DQ first divides the sampling process based on time-steps and applies a rotation matrix to smooth activations and weights dynamically. For different time-steps, a dedicated hyperparameter is introduced for adaptive timing modeling, which enables dynamic quantization across different time steps. Additionally, we also explore the compression potential of Classifier-Free Guidance (CFG-wise) to establish a foundation for subsequent work. TR-DQ achieves state-of-the-art (SOTA) performance on image generation and video generation tasks and a 1.38-1.89x speedup and 1.97-2.58x memory reduction in inference compared to existing quantization methods.

扩散模型量化生成模型推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。