arXiv:2605.00229stat.MLcs.LG2026-05被引 2

统一了扩散与流模型的微调与采样方法,揭示了不同算法的稳定性差异。

A unified perspective on fine-tuning and sampling with diffusion and flow models

论文配图:A unified perspective on fine-tuning and sampling with diffusion and flow models
图 1 · 摘自论文原文
  • 从控制论与非平衡热力学出发,构建统一分析框架。
  • 证明邻接匹配法梯度方差有限,而其他方法存在无限方差风险。
  • 适用于需要稳定微调生成模型的研究者,如可控图像生成任务。

我们研究了在指数倾斜基密度定义的目标分布上训练扩散与流生成模型的问题;该设定涵盖了从非归一化密度采样和预训练模型的奖励微调。该问题可从随机最优控制(SOC)视角,采用伴随法或得分匹配方法,或从非平衡热力学视角处理。本文提供了一个涵盖这些方法的统一框架,并有三项主要贡献:(i) 偏差-方差分解表明,伴随匹配/采样与新型得分匹配具有有限梯度方差,而目标与条件得分匹配则不具备;(ii) 对简化伴随常微分方程(ODE)给出了范数界,理论上支持伴随法的有效性;(iii) 将CMCD与NETS损失函数适配到指数倾斜设置,并提出了新的Crooks与Jarzynski恒等式。我们在Stable Diffusion 1.5和3上通过奖励微调实验验证了分析结果。

原文摘要 · Abstract (English)

We study the problem of training diffusion and flow generative models to sample from target distributions defined by an exponential tilting of a base density; a formulation that subsumes both sampling from unnormalized densities and reward fine-tuning of pre-trained models. This problem can be approached from a stochastic optimal control (SOC) perspective, using adjoint-based or score matching methods, or from a non-equilibrium thermodynamics perspective. We provide a unified framework encompassing these approaches and make three main contributions: (i) bias-variance decompositions revealing that Adjoint Matching/Sampling and Novel Score Matching have finite gradient variance, while Target and Conditional Score Matching do not; (ii) norm bounds on the lean adjoint ODE that theoretically support the effectiveness of adjoint-based methods; and (iii) adaptations of the CMCD and NETS loss functions, along with novel Crooks and Jarzynski identities, to the exponential tilting setting. We validate our analysis with reward fine-tuning experiments on Stable Diffusion 1.5 and 3.

生成模型扩散模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。