arXiv:2506.09416cs.CV2025-06ICML被引 2

将扩散模型蒸馏为快速生成的去噪器,支持多步优化与零样本控制采样。

Noise Conditional Variational Score Distillation

  • 利用噪声条件下的变分分数蒸馏,从无条件分数函数推导去噪后验分布
  • 测试时可扩展计算量,生成质量超越教师模型,逆问题上LPIPS创纪录
  • 支持单步高速生成、多步提升质量,适合需要快速可控生成的场景

我们提出噪声条件变分分数蒸馏(NCVSD),一种将预训练扩散模型蒸馏为生成性去噪器的新方法。通过揭示无条件分数函数隐式表征了去噪后验分布的分数函数,我们将此洞察融入变分分数蒸馏(VSD)框架,实现可扩展的学习,使生成性去噪器能在多种噪声水平下近似采样自去噪后验分布。所提去噪器具备良好特性:在高噪声水平下通过纯高斯噪声采样实现快速单步生成;通过增加测试时计算量实现多步采样以提升样本质量;支持零样本概率推理,实现灵活可控采样。我们在大量实验中评估NCVSD,涵盖类别条件图像生成与逆问题求解。通过扩展测试时计算量,该方法性能超越教师扩散模型,并与更大规模一致性模型相当。此外,相比基于扩散的方法,显著减少NFE(数值积分步数),在逆问题上取得突破性的LPIPS表现。

原文摘要 · Abstract (English)

We propose Noise Conditional Variational Score Distillation (NCVSD), a novel method for distilling pretrained diffusion models into generative denoisers. We achieve this by revealing that the unconditional score function implicitly characterizes the score function of denoising posterior distributions. By integrating this insight into the Variational Score Distillation (VSD) framework, we enable scalable learning of generative denoisers capable of approximating samples from the denoising posterior distribution across a wide range of noise levels. The proposed generative denoisers exhibit desirable properties that allow fast generation while preserve the benefit of iterative refinement: (1) fast one-step generation through sampling from pure Gaussian noise at high noise levels; (2) improved sample quality by scaling the test-time compute with multi-step sampling; and (3) zero-shot probabilistic inference for flexible and controllable sampling. We evaluate NCVSD through extensive experiments, including class-conditional image generation and inverse problem solving. By scaling the test-time compute, our method outperforms teacher diffusion models and is on par with consistency models of larger sizes. Additionally, with significantly fewer NFEs than diffusion-based methods, we achieve record-breaking LPIPS on inverse problems.

扩散模型模型蒸馏去噪器生成加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。