arXiv:2604.19009cs.LGcs.CV2026-04被引 2

用梯度指导强化学习,让快速生成模型更高质量。

Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning

论文配图:Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning
图 1 · 摘自论文原文
  • 以蒸馏梯度代替原始图像评分,构建更可靠的奖励信号
  • 4步生成效果超越多步教师模型,刷新少步生成性能纪录
  • 适合追求高速高质生成的落地应用开发者

扩散蒸馏(如分布匹配蒸馏,DMD)在少步生成中表现优异,但常以牺牲质量为代价换取速度。将强化学习(RL)引入蒸馏虽具潜力,但直接使用原始样本评分会导致优化目标冲突,且早期生成噪声大,奖励不可靠。为此,我们提出GDMD框架,将DMD梯度重新解释为隐式目标张量,使现有奖励模型可直接评估蒸馏更新的质量。该梯度级指导作为自适应加权机制,使强化学习策略与蒸馏目标对齐,有效消除优化偏差。实验表明,GDMD在少步生成上达到新SOTA:4步模型在GenEval和人类偏好评测中均优于多步教师模型,显著超越此前DMDR方法,展现出强可扩展性。

原文摘要 · Abstract (English)

Diffusion distillation, exemplified by Distribution Matching Distillation (DMD), has shown great promise in few-step generation but often sacrifices quality for sampling speed. While integrating Reinforcement Learning (RL) into distillation offers potential, a naive fusion of these two objectives relies on suboptimal raw sample evaluation. This sample-based scoring creates inherent conflicts with the distillation trajectory and produces unreliable rewards due to the noisy nature of early-stage generation. To overcome these limitations, we propose GDMD, a novel framework that redefines the reward mechanism by prioritizing distillation gradients over raw pixel outputs as the primary signal for optimization. By reinterpreting the DMD gradients as implicit target tensors, our framework enables existing reward models to directly evaluate the quality of distillation updates. This gradient-level guidance functions as an adaptive weighting that synchronizes the RL policy with the distillation objective, effectively neutralizing optimization divergence. Empirical results show that GDMD sets a new SOTA for few-step generation. Specifically, our 4-step models outperform the quality of their multi-step teacher and substantially exceed previous DMDR results in GenEval and human-preference metrics, exhibiting strong scalability potential.

扩散模型蒸馏强化学习少步生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。