arXiv:2605.18736cs.CV2026-05被引 3

通过频域渐进生成,实现图像视频模型的高效加速。

Spectral Progressive Diffusion for Efficient Image and Video Generation

论文配图:Spectral Progressive Diffusion for Efficient Image and Video Generation
图 1 · 摘自论文原文
  • 在去噪过程中逐步提升分辨率,利用频谱特性避免冗余计算。
  • 相比原模型,生成速度显著提升,且保持高质量视觉效果。
  • 无需重新训练,适用于主流扩散模型,适合追求效率的研究者。

扩散模型在去噪过程中隐式地以自回归方式在频域生成视觉内容,低频成分先于高频细节出现。这一结构为高效生成提供了天然机会,因为噪声主导的高频区域进行高分辨率计算是冗余的。我们提出频谱渐进扩散(Spectral Progressive Diffusion),一种通用框架,在预训练扩散模型的去噪轨迹中逐步增长分辨率。为此,我们设计了频谱噪声扩展机制,并从模型功率谱中推导出最优分辨率调度。该框架支持无训练加速和一种新的微调方法,进一步提升效率与质量。我们在最先进的预训练图像与视频生成模型上实现了显著提速,同时保持视觉保真度。

原文摘要 · Abstract (English)

Diffusion models have been shown to implicitly generate visual content autoregressively in the frequency domain, where low-frequency components are generated earlier in the denoising process while high-frequency details emerge only in later timesteps. This structure offers a natural opportunity for efficient generation, as high-resolution computation on noise-dominated frequencies is largely redundant. We propose Spectral Progressive Diffusion, a general framework that progressively grows resolution along the denoising trajectory of pretrained diffusion models. To this end, we develop a spectral noise expansion mechanism and derive an optimal resolution schedule from the model's power spectrum. Our framework supports training-free acceleration and a novel fine-tuning recipe that further improves efficiency and quality. We demonstrate significant speedups on state-of-the-art pretrained image and video generation models while preserving visual quality.

扩散模型高效生成频域建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。