通过噪声注入提升单步生成质量,实现高效多模态图像生成
RMFlow: Refined Mean Flow by a Noise-Injection Step for Multimodal Generation
- 用神经网络拟合流路径平均速度,结合新损失函数优化
- 仅用1次前向计算即达接近顶尖的生成效果
- 适合需要快速生成且质量要求高的多模态应用
均值流(MeanFlow)实现了高效的高保真图像生成,但其单一阶段评估(1-NFE)常难以生成令人满意的输出。为此,我们提出RMFlow,一种高效的多模态生成模型,将粗粒度的1-NFE均值流传输与后续定制化的噪声注入精炼步骤相结合。RMFlow利用神经网络近似流路径的平均速度,通过新设计的损失函数训练,该损失函数在最小化概率路径间的Wasserstein距离和最大化样本似然之间取得平衡。在文本到图像、上下文到分子、时间序列生成任务上,RMFlow仅使用1次前向传播即达到接近当前最优的效果,计算成本与基线均值流相当。
原文摘要 · Abstract (English)
Mean flow (MeanFlow) enables efficient, high-fidelity image generation, yet its single-function evaluation (1-NFE) generation often cannot yield compelling results. We address this issue by introducing RMFlow, an efficient multimodal generative model that integrates a coarse 1-NFE MeanFlow transport with a subsequent tailored noise-injection refinement step. RMFlow approximates the average velocity of the flow path using a neural network trained with a new loss function that balances minimizing the Wasserstein distance between probability paths and maximizing sample likelihood. RMFlow achieves near state-of-the-art results on text-to-image, context-to-molecule, and time-series generation using only 1-NFE, at a computational cost comparable to the baseline MeanFlows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。