arXiv:2512.21829stat.MLcs.LG2025-12被引 17

一种无需梯度的高效采样与微调方法,提升生成模型性能。

Tilt Matching for Scalable Sampling and Fine-Tuning

  • 基于随机插值的动态方程设计新速度场,实现无梯度优化。
  • 在Lennard-Jones势下采样效率达当前最优,微调Stable Diffusion表现良好。
  • 适用于多步流映射模型,无需奖励系数调节,适合实际部署。

我们提出一种简单且可扩展的算法,利用随机插值对未归一化密度进行采样,并用于生成模型的微调。该方法称为倾斜匹配(Tilt Matching),源于一个将流匹配速度与受奖励倾斜分布相关联的动力学方程,隐式求解了一个随机最优控制问题。新速度场继承了随机插值传输的光滑性,同时是最小化目标函数的解,其方差严格低于流匹配本身。速度场更新可解释为随机插值与奖励副本的所有联合累积量之和,一阶近似为其协方差。该算法无需访问奖励梯度,也无需对流或扩散轨迹进行反向传播。实验验证该方法高效且高度可扩展,在Lennard-Jones势下的采样任务中达到当前最优结果,且在微调Stable Diffusion上表现竞争力,无需奖励乘子。该方法还可直接应用于少步流映射模型。

原文摘要 · Abstract (English)

We propose a simple, scalable algorithm for using stochastic interpolants to sample from unnormalized densities and for fine-tuning generative models. The approach, Tilt Matching, arises from a dynamical equation relating the flow matching velocity to one targeting the same distribution tilted by a reward, implicitly solving a stochastic optimal control problem. The new velocity inherits the regularity of stochastic interpolant transports while also being the minimizer of an objective with strictly lower variance than flow matching itself. The update to the velocity field can be interpreted as the sum of all joint cumulants of the stochastic interpolant and copies of the reward, and to first order is their covariance. The algorithms do not require any access to gradients of the reward or backpropagating through trajectories of the flow or diffusion. We empirically verify that the approach is efficient and highly scalable, providing state-of-the-art results on sampling under Lennard-Jones potentials and is competitive on fine-tuning Stable Diffusion, without requiring reward multipliers. It can also be straightforwardly applied to tilting few-step flow map models.

生成模型采样算法无梯度优化扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。