将蒸馏与强化学习结合,让少步生成模型更懂人类偏好且更稳定。
Distribution Matching Distillation Meets Reinforcement Learning
- 先用动态蒸馏优化分布匹配,再联合强化学习提升可控性。
- 在少步生成中达到顶尖视觉质量,甚至超过多步教师模型。
- 适合追求高效、高可控生成的开发者和研究者使用。
分布匹配蒸馏(DMD)通过将多步扩散模型压缩为少步版本,实现高效推理。同时,强化学习(RL)已成为对齐生成模型与人类偏好的关键工具。尽管两者都是大规模扩散模型训练后的关键阶段,现有研究通常将其视为独立、顺序的过程,缺乏系统性统一框架。本文表明,联合优化二者可互惠共赢:强化学习使蒸馏更具偏好感知和可控性,而非均匀压缩数据分布;而DMD则有效缓解强化学习中的奖励滥用问题。基于此,我们提出DMDR框架,在第一阶段采用带奖励倾斜的分布匹配与两种动态蒸馏策略,第二阶段进行联合DMD与RL优化。大量实验表明,DMDR在少步生成中达到当前最优的视觉质量和提示遵循度,甚至超越其多步教师模型。
原文摘要 · Abstract (English)
Distribution Matching Distillation (DMD) facilitates efficient inference by distilling multi-step diffusion models into few-step variants. Concurrently, Reinforcement Learning (RL) has emerged as a vital tool for aligning generative models with human preferences. While both represent critical post-training stages for large-scale diffusion models, existing studies typically treat them as independent, sequential processes, leaving a systematic framework for their unification largely unexplored. In this work, we demonstrate that jointly optimizing these two objectives yields mutual benefits: RL enables more preference-aware and controllable distillation rather than uniformly compressing the full data distribution, while DMD serves as an effective regularizer to mitigate reward hacking during RL training. Building on these insights, we propose DMDR, a unified framework that incorporates Reward-Tilted Distribution Matching optimization alongside two dynamic distillation training strategies in the initial stage, followed by the joint DMD and RL optimization in the second stage. Extensive experiments demonstrate that DMDR achieves state-of-the-art visual quality and prompt adherence among few-step generation methods, even surpassing the performance of its multi-step teacher model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。