用流映射提升扩散模型测试时采样质量,让生成图像更符合用户奖励。
Test-time scaling of diffusions with flow maps
- 基于流映射与速度场关系,设计新算法FMTT优化奖励梯度方向。
- 在多个任务中优于传统方法,能准确找到奖励分布的局部最优解。
- 适合需要精细控制生成结果的场景,如与视觉语言模型联动编辑图像。
为提升扩散模型在测试阶段生成样本的用户指定奖励得分,通常将奖励梯度引入扩散过程。但此类方法常因奖励仅在生成末端的数据分布上定义良好而病态。现有方法多依赖去噪器估计末态样本,我们提出直接使用流映射的简单方案。利用流映射与瞬时传输速度场之间的关系,构建了流动轨迹倾斜算法(FMTT),其在奖励上升性能上可证明优于标准测试时方法。该方法可用于精确采样(重要性加权)或有原则的搜索,以识别奖励倾斜分布的局部最大值。实验表明,相比其他前瞻技术,本方法表现更优,并使复杂奖励函数成为可能,例如通过对接视觉语言模型实现新型图像编辑。
原文摘要 · Abstract (English)
A common recipe to improve diffusion models at test-time so that samples score highly against a user-specified reward is to introduce the gradient of the reward into the dynamics of the diffusion itself. This procedure is often ill posed, as user-specified rewards are usually only well defined on the data distribution at the end of generation. While common workarounds to this problem are to use a denoiser to estimate what a sample would have been at the end of generation, we propose a simple solution to this problem by working directly with a flow map. By exploiting a relationship between the flow map and velocity field governing the instantaneous transport, we construct an algorithm, Flow Map Trajectory Tilting (FMTT), which provably performs better ascent on the reward than standard test-time methods involving the gradient of the reward. The approach can be used to either perform exact sampling via importance weighting or principled search that identifies local maximizers of the reward-tilted distribution. We demonstrate the efficacy of our approach against other look-ahead techniques, and show how the flow map enables engagement with complicated reward functions that make possible new forms of image editing, e.g. by interfacing with vision language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。