arXiv:2603.06136cs.CV2026-03被引 1

解决跨分辨率生成质量下降问题,实现快速高保真图像生成

Cross-Resolution Distribution Matching for Diffusion Distillation

  • 用对数信噪比曲线划分不同分辨率的推理步骤,精准匹配分布
  • 在SDXL上实现33.4倍加速,保持高质量图像输出
  • 适合追求高效生成且注重画质的研究者与开发者

扩散蒸馏是加速图像与视频生成的核心技术,但现有方法受限于去噪过程,步数减少已趋于饱和。部分时间步低分辨率生成虽可进一步提速,却因跨分辨率分布差异导致明显质量下降。本文提出跨分辨率分布匹配蒸馏(RMD),一种新框架,通过在各分辨率下基于对数信噪比(logSNR)曲线划分时间步,并引入logSNR映射补偿分辨率带来的分布偏移。沿分辨率轨迹进行分布匹配,缩小低分辨率生成器分布与教师模型高分辨率分布之间的差距。同时,在上采样阶段引入预测噪声重注入机制,稳定训练并提升合成质量。定量与定性结果表明,RMD在多种骨干网络上均实现高保真、少步数多分辨率级联推理。特别地,RMD在SDXL上达到33.4倍加速,在Wan2.1-14B上实现25.6倍加速,同时保持高视觉保真度。

原文摘要 · Abstract (English)

Diffusion distillation is central to accelerating image and video generation, yet existing methods are fundamentally limited by the denoising process, where step reduction has largely saturated. Partial timestep low-resolution generation can further accelerate inference, but it suffers noticeable quality degradation due to cross-resolution distribution gaps. We propose Cross-Resolution Distribution Matching Distillation (RMD), a novel distillation framework that bridges cross-resolution distribution gaps for high-fidelity, few-step multi-resolution cascaded inference. Specifically, RMD divides the timestep intervals for each resolution using logarithmic signal-to-noise ratio (logSNR) curves, and introduces logSNR-based mapping to compensate for resolution-induced shifts. Distribution matching is conducted along resolution trajectories to reduce the gap between low-resolution generator distributions and the teacher's high-resolution distribution. In addition, a predicted-noise re-injection mechanism is incorporated during upsampling to stabilize training and improve synthesis quality. Quantitative and qualitative results show that RMD preserves high-fidelity generation while accelerating inference across various backbones. Notably, RMD achieves up to 33.4X speedup on SDXL and 25.6X on Wan2.1-14B, while preserving high visual fidelity.

扩散模型加速生成跨分辨率蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。