arXiv:2606.00658cs.CVcs.AI2026-06

压缩视频扩散模型,用少步推理+低比特量化提升效率

Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models

论文配图:Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models
图 1 · 摘自论文原文
  • 分路校准高低噪声分支,保护关键层,优化低比特表示
  • 20步推理时质量超越原模型,8步和20步平均表现更优
  • 适合需要高效部署的视频生成场景,尤其关注推理速度

大型视频扩散模型虽视觉质量优异,但部署成本高,因每样本需大量去噪步骤且参数量庞大。本文针对Wan2.2-T2V-A14B提出面向部署的压缩流水线,结合少步分布匹配蒸馏与低比特量化。该流程遵循双专家去噪路径,分别校准高噪声与低噪声分支,保护敏感输入层,并采用HiF4风格低比特表示以增强动态范围覆盖。量化在蒸馏后的少步学生模型上进行,而非原始长步轨迹,降低推理时激活分布不匹配。所提协同设计使量化模型在相同步数下接近全精度模型,且在8步和20步平均表现超越原全精度基线。20步设置在测试配置中实现最佳质量-效率平衡。

原文摘要 · Abstract (English)

Large video diffusion models achieve strong visual quality but remain expensive to deploy because each sample requires many denoising steps and a large resident parameter footprint. This paper studies a deployment-oriented compression pipeline for Wan2.2-T2V-A14B by combining few-step distribution-matching distillation with low-bit quantization. The pipeline follows the model's dual-expert denoising route, calibrates the high-noise and low-noise branches separately, protects sensitive entrance layers, and uses HiF4-style low-bit representation to improve dynamic-range coverage. Quantization is calibrated on the distilled few-step student rather than on the original long-step trajectory, reducing activation-distribution mismatch during inference. The proposed co-design keeps the quantized model close to the same-step full-precision model and surpasses the original full-precision baseline at 8 and 20 steps on average. The 20-step setting gives the best quality-efficiency trade-off in the tested configurations.

视频生成扩散模型量化少步推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。