arXiv:2602.00849cs.LGcs.AI2026-02中稿 · ICLR被引 5

通过噪声注入提升单步生成质量,实现高效多模态图像生成

RMFlow: Refined Mean Flow by a Noise-Injection Step for Multimodal Generation

  • 用神经网络拟合流路径平均速度,结合新损失函数优化
  • 仅用1次前向计算即达接近顶尖的生成效果
  • 适合需要快速生成且质量要求高的多模态应用

均值流(MeanFlow)实现了高效的高保真图像生成,但其单一阶段评估(1-NFE)常难以生成令人满意的输出。为此,我们提出RMFlow,一种高效的多模态生成模型,将粗粒度的1-NFE均值流传输与后续定制化的噪声注入精炼步骤相结合。RMFlow利用神经网络近似流路径的平均速度,通过新设计的损失函数训练,该损失函数在最小化概率路径间的Wasserstein距离和最大化样本似然之间取得平衡。在文本到图像、上下文到分子、时间序列生成任务上,RMFlow仅使用1次前向传播即达到接近当前最优的效果,计算成本与基线均值流相当。

原文摘要 · Abstract (English)

Mean flow (MeanFlow) enables efficient, high-fidelity image generation, yet its single-function evaluation (1-NFE) generation often cannot yield compelling results. We address this issue by introducing RMFlow, an efficient multimodal generative model that integrates a coarse 1-NFE MeanFlow transport with a subsequent tailored noise-injection refinement step. RMFlow approximates the average velocity of the flow path using a neural network trained with a new loss function that balances minimizing the Wasserstein distance between probability paths and maximizing sample likelihood. RMFlow achieves near state-of-the-art results on text-to-image, context-to-molecule, and time-series generation using only 1-NFE, at a computational cost comparable to the baseline MeanFlows.

生成模型多模态高效生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。