arXiv:2608.29647cs.LG2026-08

用最优传输方法提升单步生成模型的奖励调优效果

Reward-guided Fine-Tuning of One-Step Generative Models via Wasserstein Gradient Flow

论文配图:Reward-guided Fine-Tuning of One-Step Generative Models via Wasserstein Gradient Flow
图 1 · 摘自论文原文
  • 基于最优传输视角,利用Wasserstein梯度流实现平滑分布演化
  • 无需奖励梯度即可稳定更新,有效避免奖励欺骗和模式崩溃
  • 在多种图像数据上表现优于基线,支持复杂奖励函数

为降低生成模型的时间复杂度,单步生成模型通过一次前向传播直接将噪声映射到数据。然而,此类模型的奖励引导微调方法仍不成熟。本文从最优传输角度出发,研究用于概率空间中平滑可控分布演化的Wasserstein梯度流(WGF),并提出一种基于WGF的新型单步生成模型奖励引导微调方法。该方法无需奖励梯度,可处理可导与不可导奖励,实现稳定平滑的分布更新,有效缓解奖励欺骗和模式崩溃问题。在2D合成数据、CIFAR-10及ImageNet 256×256上的实验表明,使用多种奖励(包括JPEG压缩性、类别概率、黑白化、CLIP对齐)时,本方法在奖励对齐方面优于现有基线。

原文摘要 · Abstract (English)

To mitigate the time complexity of generative models, one-step generative models have recently emerged through direct mapping from noise to data in a single forward pass. However, the reward-guided fine-tuning method of one-step generative models remains largely unexplored. To address this, we consider one-step generators from an optimal transport view, investigating Wasserstein Gradient Flow (WGF) for modeling smooth and controlled distributional evolution in probability space. We then propose a novel reward-guided fine-tuning of a one-step generative model via WGF. We derive a practical training method that requires no reward gradients, thereby handling both non-differentiable and differentiable rewards. Moreover, our method provides smooth and stable reward-guided distributional updates while mitigating reward hacking and mode collapse. Experiments on 2D synthetic data, CIFAR-10, and ImageNet 256$\times$256 with diverse rewards, including JPEG (in)compressibility, class probability, Black-and-White and CLIP alignment, show that our method achieves better reward alignment compared to baselines.

生成模型最优传输奖励调优扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。