arXiv:2606.30414cs.LG2026-06被引 2

将强化学习与蒸馏结合,提升扩散模型生成质量与速度。

Diffusion Fine-tuning with Rewarded Moment Matching Distillation

论文配图:Diffusion Fine-tuning with Rewarded Moment Matching Distillation
图 1 · 摘自论文原文
  • 用奖励机制优化蒸馏过程,让模型生成更自然的图像。
  • 在ImageNet上实现更高精度与更快推理的平衡,优于现有方法。
  • 适用于气象等复杂科学领域,可显著提速且性能更优。

蒸馏与强化学习微调是扩散模型后训练的主要支柱。尽管两者常被独立研究,其交互关系仍不明确,尤其是微调如何影响蒸馏模型的生成质量。本文提出奖励时刻匹配蒸馏(RMMD),一种同时蒸馏扩散模型并最大化奖励函数的新框架。RMMD通过调整采样循环实现在线策略训练,并将蒸馏损失重新用作积分KL正则化的代理,从而保持先进蒸馏技术(如8步时刻匹配)的高保真度‘自然性’特征。在ImageNet上评估FID-奖励帕累托前沿,结果表明RMMD在对比单步基线(DI++)和多步竞争方法(DRaFT、HyperNoise)时表现出更优的权衡。此外,我们将RMMD应用于最先进的天气预测模型GenCast,蒸馏过程中优化连续排名概率评分(CRPS)。所得蒸馏模型实现7.5倍加速,且在93%的目标气象变量上表现优于教师模型,校准性更优。这证明了RMMD可扩展至复杂高维科学领域。

原文摘要 · Abstract (English)

Distillation and Reinforcement Learning (RL) fine-tuning are the primary pillars of diffusion post-training. While traditionally studied in isolation, the interaction between these phases remains poorly understood, and in particular how fine-tuning impacts the generative quality of distilled models. We introduce Rewarded Moment Matching Distillation (RMMD), a novel framework that simultaneously distills diffusion models and maximizes a reward function. RMMD preserves the high-fidelity ``naturalness'' characteristic of advanced distillation (such as 8-step Moment Matching) by adapting the sampling loop for on-policy training and repurposing the distillation loss as a proxy for integral KL regularization. By evaluating the FID-Reward Pareto fronts on ImageNet, we demonstrate that RMMD achieves superior trade-offs compared to single-step baselines (DI++) and multi-step competitors (DRaFT, HyperNoise). Finally, we apply RMMD to GenCast, a state-of-the-art weather forecasting model, to distill it while optimizing the Continuous Ranked Probability Score (CRPS) metric. The resulting distilled model achieves a 7.5x speedup while outperforming the teacher model on 93% of target weather variables, and being better calibrated. This proves that RMMD scales to complex, high-dimensional scientific domains.

扩散模型蒸馏强化学习气象预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。