arXiv:2507.04290cs.CV2025-07TPAMI被引 6

2-4比特下仍能保持生成质量的低比特扩散模型量化新方法

MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation

论文配图:MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation
图 1 · 摘自论文原文
  • 分层残差量化+自适应步长,精准处理关键误差
  • 在2-4比特下生成图像质量显著优于现有方法
  • 适合边缘设备部署的低比特扩散模型应用

扩散模型在视觉生成任务中表现卓越,但高计算复杂度限制了其在边缘设备的应用。量化成为加速推理和降低内存消耗的有前景技术。然而,现有量化方法在极低比特(2-4比特)下泛化能力差,直接应用会导致严重性能下降。我们发现现有框架存在异常值不友好量化器设计、次优初始化与优化策略问题。提出MPQ-DMv2,一种针对极低比特扩散模型的改进混合精度量化框架。从量化角度,针对显著异常值导致的分布失衡,提出灵活Z阶残差混合量化,利用高效二进制残差分支实现可调量化步长以处理关键误差。从优化框架看,理论上分析了LoRA模块的收敛性与最优性,提出面向对象的低秩初始化,利用先验量化误差进行信息初始化。进一步提出基于记忆的时序关系蒸馏,构建在线时间感知像素队列,实现长期去噪时序信息的蒸馏,保障量化模型与全精度模型之间的整体时序一致性。在多种生成任务上的全面实验表明,MPQ-DMv2在不同架构上均显著超越当前最先进方法,尤其在极低比特宽度下优势明显。

原文摘要 · Abstract (English)

Diffusion models have demonstrated remarkable performance on vision generation tasks. However, the high computational complexity hinders its wide application on edge devices. Quantization has emerged as a promising technique for inference acceleration and memory reduction. However, existing quantization methods do not generalize well under extremely low-bit (2-4 bit) quantization. Directly applying these methods will cause severe performance degradation. We identify that the existing quantization framework suffers from the outlier-unfriendly quantizer design, suboptimal initialization, and optimization strategy. We present MPQ-DMv2, an improved \textbf{M}ixed \textbf{P}recision \textbf{Q}uantization framework for extremely low-bit \textbf{D}iffusion \textbf{M}odels. For the quantization perspective, the imbalanced distribution caused by salient outliers is quantization-unfriendly for uniform quantizer. We propose \textit{Flexible Z-Order Residual Mixed Quantization} that utilizes an efficient binary residual branch for flexible quant steps to handle salient error. For the optimization framework, we theoretically analyzed the convergence and optimality of the LoRA module and propose \textit{Object-Oriented Low-Rank Initialization} to use prior quantization error for informative initialization. We then propose \textit{Memory-based Temporal Relation Distillation} to construct an online time-aware pixel queue for long-term denoising temporal information distillation, which ensures the overall temporal consistency between quantized and full-precision model. Comprehensive experiments on various generation tasks show that our MPQ-DMv2 surpasses current SOTA methods by a great margin on different architectures, especially under extremely low-bit widths.

扩散模型量化低比特边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。