arXiv:2510.02692cs.LGcs.AI2025-10中稿 · ICLR被引 3

通过调整中间噪声分布,提升扩散模型微调效果

Fine-Tuning Diffusion Models via Intermediate Distribution Shaping

  • 在中间噪声水平上重塑分布,实现更优微调
  • 文本到图像生成中相对基线提升8.81%的VQAScore
  • 无需显式奖励即可修正预训练流模型误差

扩散模型广泛应用于各类生成任务。给定预训练扩散模型时,常需进一步微调以纠正学习误差或适配下游应用。本文研究扩散模型在中间噪声水平诱导分布的影响。首先,我们统一现有基于拒绝采样的微调方法(RAFT),提出GRAFT,并揭示其等价于带重分布奖励的KL正则化奖励最大化。受此启发,提出P-GRAFT,在中间噪声水平上主动塑造分布,实证表明可提升微调效率,并通过偏差-方差权衡提供数学解释。进一步基于该框架提出逆噪声校正算法,无需显式奖励即可改善预训练流模型质量。在文本到图像生成、布局生成、分子生成和无条件图像生成任务上评估,应用于Stable Diffusion v2时,在主流文本到图像基准上优于策略梯度方法,VQAScore相对提升8.81%;在无条件图像生成中,以更低每图像浮点运算量(FLOPs/image)改善生成图像的FID得分。

原文摘要 · Abstract (English)

Diffusion models are widely used for generative tasks across domains. Given a pre-trained diffusion model, it is often desirable to fine-tune it further either to correct for errors in learning or to align with downstream applications. Towards this, we examine the effect of shaping the distribution at intermediate noise levels induced by diffusion models. First, we show that existing variants of Rejection sAmpling based Fine-Tuning (RAFT), which we unify as GRAFT, can implicitly perform KL regularized reward maximization with reshaped rewards. Motivated by this observation, we introduce P-GRAFT to shape distributions at intermediate noise levels and demonstrate empirically that this can lead to more effective fine-tuning. We mathematically explain this via a bias-variance tradeoff. Next, we look at correcting learning errors in pre-trained flow models based on the developed mathematical framework. In particular, we propose inverse noise correction, a novel algorithm to improve the quality of pre-trained flow models without explicit rewards. We empirically evaluate our methods on text-to-image(T2I) generation, layout generation, molecule generation and unconditional image generation. Notably, our framework, applied to Stable Diffusion v2, improves over policy gradient methods on popular T2I benchmarks in terms of VQAScore and shows an $8.81\%$ relative improvement over the base model. For unconditional image generation, inverse noise correction improves FID of generated images at lower FLOPs/image.

扩散模型微调分布重塑生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。