根据图像频谱特性设计噪声调度,提升生成质量。
Spectrally-Guided Diffusion Noise Schedules
- 基于图像频谱特性自动设计噪声调度,避免人工调参。
- 在低采样步数下显著提升单阶段像素扩散模型的生成质量。
- 适合追求高效高质图像生成的研究者和开发者。
去噪扩散模型广泛用于高质量图像和视频生成。其性能依赖于噪声调度,即训练时施加的噪声水平分布及采样过程中的噪声序列。传统噪声调度通常为手工设计,需针对不同分辨率进行手动调优。本文提出一种基于图像频谱特性的实例化噪声调度设计方法。通过推导最小与最大噪声水平的有效性理论边界,我们设计出“紧致”噪声调度,消除了冗余步骤。推理时,我们提出条件采样此类噪声调度。实验表明,该方法显著提升了单阶段像素扩散模型的生成质量,尤其在低步数条件下表现优异。
原文摘要 · Abstract (English)
Denoising diffusion models are widely used for high-quality image and video generation. Their performance depends on noise schedules, which define the distribution of noise levels applied during training and the sequence of noise levels traversed during sampling. Noise schedules are typically handcrafted and require manual tuning across different resolutions. In this work, we propose a principled way to design per-instance noise schedules for pixel diffusion, based on the image's spectral properties. By deriving theoretical bounds on the efficacy of minimum and maximum noise levels, we design ``tight'' noise schedules that eliminate redundant steps. During inference, we propose to conditionally sample such noise schedules. Experiments show that our noise schedules improve generative quality of single-stage pixel diffusion models, particularly in the low-step regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。