改进生成模型轨迹,实现快速高质量图像生成
Straighten Viscous Rectified Flow via Noise Optimization
- 引入历史速度项与噪声优化,构建更准确的生成路径
- 在单步和少步生成中达到当前最佳性能
- 适合需要高效高质图像生成的研究与应用
Reflow 操作旨在通过构建噪声与图像间的确定性耦合,在训练中拉直修正流的推理轨迹,从而提升单步或少步生成图像的质量。然而,我们发现 Reflow 存在关键缺陷:其构造的确定性耦合中图像分布与真实图像存在差距,导致难以快速生成高质量图像。为此,我们提出一种新方法——通过噪声优化拉直黏性修正流(VRFNO),这是一个联合训练框架,包含编码器与神经速度场。VRFNO 的两大创新为:(1) 引入历史速度项以增强轨迹区分度,使模型更准确预测当前轨迹速度;(2) 通过重参数化进行噪声优化,形成与真实图像匹配的优化耦合,用于训练,有效缓解 Reflow 的误差问题。在合成数据及多种分辨率的真实数据集上的全面实验表明,VRFNO 显著缓解了 Reflow 的局限性,在单步与少步生成任务中均达到最先进水平。
原文摘要 · Abstract (English)
The Reflow operation aims to straighten the inference trajectories of the rectified flow during training by constructing deterministic couplings between noises and images, thereby improving the quality of generated images in single-step or few-step generation. However, we identify critical limitations in Reflow, particularly its inability to rapidly generate high-quality images due to a distribution gap between images in its constructed deterministic couplings and real images. To address these shortcomings, we propose a novel alternative called Straighten Viscous Rectified Flow via Noise Optimization (VRFNO), which is a joint training framework integrating an encoder and a neural velocity field. VRFNO introduces two key innovations: (1) a historical velocity term that enhances trajectory distinction, enabling the model to more accurately predict the velocity of the current trajectory, and (2) the noise optimization through reparameterization to form optimized couplings with real images which are then utilized for training, effectively mitigating errors caused by Reflow's limitations. Comprehensive experiments on synthetic data and real datasets with varying resolutions show that VRFNO significantly mitigates the limitations of Reflow, achieving state-of-the-art performance in both one-step and few-step generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。