arXiv:2605.27003cs.CVcs.AI2026-05被引 1

针对视频扩散模型的4位量化难题,提出分时步、分专家的精准校准方法。

Timestep-Aware SVDQuant-GPTQ for W4A4 Quantization of Wan2.2-I2V

论文配图:Timestep-Aware SVDQuant-GPTQ for W4A4 Quantization of Wan2.2-I2V
图 1 · 摘自论文原文
  • 分时步、分专家设计独立量化策略,应对激活分布差异
  • 峰值显存降低59.3%,视觉质量仅下降0.9%(VBench)
  • 适合追求高保真4位量化推理的视频生成研究者

大视频扩散模型的W4A4量化虽能大幅节省内存,但面临两大挑战:大数值激活异常值稀疏分布,以及多步去噪轨迹中时序依赖的激活分布差异。这一问题在双专家混合专家结构(MoE DiT)的Wan2.2-I2V模型中尤为突出,其高噪声与低噪声专家对量化敏感度不同,单一全局校准策略难以兼顾。本文提出后训练量化框架,结合基于SVDQuant的低秩异常值补偿、GPTQ驱动的重建感知残差权重量化,以及按时步分箱、每层独立进行的专家级激活截断比搜索。在OpenS2V-Eval基准上,相较BF16基线,本方法将峰值显存降低59.3%,仅导致VBench平均得分下降0.9%、成像质量下降2.3%,证明针对专家与时步的感知校准对高保真W4A4推理至关重要。

原文摘要 · Abstract (English)

W4A4 quantization of large video diffusion Transformers offers substantial memory savings but is hindered by two main challenges: sparse large-magnitude activation outliers, and strongly timestep-dependent activation distributions across the multi-step denoising trajectory. These difficulties are compounded by Wan2.2-I2V's two-expert Mixture-of-Experts DiT design, whose high-noise and low-noise experts exhibit distinct quantization sensitivities that a single global calibration policy cannot capture. We propose a post-training quantization framework combining SVDQuant-based low-rank outlier compensation, GPTQ-based reconstruction-aware residual weight quantization, and timestep-bin-wise per-layer activation clipping-ratio search conducted independently for each expert. On the OpenS2V-Eval benchmark, our method reduces peak GPU memory by 59.3\% relative to the BF16 baseline while incurring only a 0.9\% drop in VBench average score and a 2.3\% drop in Imaging Quality, demonstrating that expert- and timestep-aware calibration is essential for high-fidelity W4A4 inference on MoE video DiTs.

视频生成量化MoE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。