arXiv:2503.06652cs.CV2025-03ICCV被引 10

让一步生成模型轻松适配新控制条件,无需重训。

Adding Additional Control to One-Step Diffusion with Joint Distribution Matching

  • 通过联合分布匹配,分离图像保真与条件学习任务。
  • 仅用一步生成就超越多步ControlNet,且支持新控制条件。
  • 适合需要快速迭代控制逻辑或融合用户反馈的场景。

尽管扩散蒸馏已实现一步生成,但将蒸馏模型适配到新控制条件(如新结构约束或用户偏好)仍具挑战。传统方法需修改基础扩散模型并重新蒸馏,计算成本高、耗时长。为此,我们提出联合分布匹配(JDM),最小化图像-条件联合分布的反向KL散度。通过推导可计算上界,JDM实现了保真度学习与条件学习的解耦。该非对称蒸馏方案使一步学生模型能处理教师模型未见过的控制条件,提升无分类器引导(CFG)使用效率,并支持人机反馈学习(HFL)无缝集成。实验表明,JDM在多数情况下仅用一步生成即超越多步ControlNet,且在一步文本到图像生成中达到当前最优性能,得益于更优的CFG利用或HFL整合。

原文摘要 · Abstract (English)

While diffusion distillation has enabled one-step generation through methods like Variational Score Distillation, adapting distilled models to emerging new controls -- such as novel structural constraints or latest user preferences -- remains challenging. Conventional approaches typically requires modifying the base diffusion model and redistilling it -- a process that is both computationally intensive and time-consuming. To address these challenges, we introduce Joint Distribution Matching (JDM), a novel approach that minimizes the reverse KL divergence between image-condition joint distributions. By deriving a tractable upper bound, JDM decouples fidelity learning from condition learning. This asymmetric distillation scheme enables our one-step student to handle controls unknown to the teacher model and facilitates improved classifier-free guidance (CFG) usage and seamless integration of human feedback learning (HFL). Experimental results demonstrate that JDM surpasses baseline methods such as multi-step ControlNet by mere one-step in most cases, while achieving state-of-the-art performance in one-step text-to-image synthesis through improved usage of CFG or HFL integration.

扩散模型一步生成控制条件分布匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。