让生成模型学会多方向运动,提升图像生成质量。
Variational Rectified Flow Matching
- 用变分方法建模多模态速度场,避免平均化真实方向。
- 在MNIST、CIFAR-10和ImageNet上生成效果显著提升。
- 适合追求高质量生成结果的扩散模型研究者。
我们研究了变分修正流匹配(Variational Rectified Flow Matching),该框架通过建模多模态速度场来增强经典修正流匹配。经典方法在推理时通过求解常微分方程,沿速度场将样本从源分布移动到目标分布;训练时通过线性插值源与目标分布的配对样本,得到指向不同方向的“真实”速度场,即多模态/模糊速度场。但因使用标准均方误差损失,学习到的速度场会平均真实方向,失去多模态特性。而变分修正流匹配则能学习并采样多模态流方向。在合成数据、MNIST、CIFAR-10和ImageNet上的实验表明,该方法取得了优异结果。
原文摘要 · Abstract (English)
We study Variational Rectified Flow Matching, a framework that enhances classic rectified flow matching by modeling multi-modal velocity vector-fields. At inference time, classic rectified flow matching 'moves' samples from a source distribution to the target distribution by solving an ordinary differential equation via integration along a velocity vector-field. At training time, the velocity vector-field is learnt by linearly interpolating between coupled samples one drawn from the source and one drawn from the target distribution randomly. This leads to ''ground-truth'' velocity vector-fields that point in different directions at the same location, i.e., the velocity vector-fields are multi-modal/ambiguous. However, since training uses a standard mean-squared-error loss, the learnt velocity vector-field averages ''ground-truth'' directions and isn't multi-modal. In contrast, variational rectified flow matching learns and samples from multi-modal flow directions. We show on synthetic data, MNIST, CIFAR-10, and ImageNet that variational rectified flow matching leads to compelling results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。