arXiv:2512.02012cs.CVcs.LG2025-12被引 88

改进快速生成模型的训练与引导机制,实现单次计算高质图像生成。

Improved Mean Flows: On the Challenges of Fastforward Generative Models

  • 重构训练目标为速度回归问题,提升训练稳定性。
  • 单次前向计算(1-NFE)下在ImageNet上达1.72 FID,性能领先。
  • 支持灵活条件控制,适合追求高效生成的研究者使用。

MeanFlow(MF)作为一种单步生成建模框架已被确立,但其‘快速前进’特性带来了训练目标与引导机制的关键挑战。首先,原始MF的训练目标不仅依赖真实场,还受网络自身影响。为此,我们重新定义目标为对瞬时速度 $v$ 的损失,通过预测平均速度 $u$ 的网络进行参数化,使问题更接近标准回归,提升了训练稳定性。其次,原始MF在训练中固定无分类器引导尺度,牺牲了灵活性。我们通过将引导显式建模为条件变量,保留测试时的灵活性。多样条件通过上下文条件处理,减少模型规模并提升性能。整体而言,我们的改进版MeanFlow(iMF)从零训练,在ImageNet 256×256上实现1.72 FID,仅需一次函数评估(1-NFE)。iMF显著优于同类单步方法,并在不使用蒸馏的情况下缩小了与多步方法的差距。我们希望本工作能推动快速前进生成建模作为独立范式的进一步发展。

原文摘要 · Abstract (English)

MeanFlow (MF) has recently been established as a framework for one-step generative modeling. However, its ``fastforward'' nature introduces key challenges in both the training objective and the guidance mechanism. First, the original MF's training target depends not only on the underlying ground-truth fields but also on the network itself. To address this issue, we recast the objective as a loss on the instantaneous velocity $v$, re-parameterized by a network that predicts the average velocity $u$. Our reformulation yields a more standard regression problem and improves the training stability. Second, the original MF fixes the classifier-free guidance scale during training, which sacrifices flexibility. We tackle this issue by formulating guidance as explicit conditioning variables, thereby retaining flexibility at test time. The diverse conditions are processed through in-context conditioning, which reduces model size and benefits performance. Overall, our $\textbf{improved MeanFlow}$ ($\textbf{iMF}$) method, trained entirely from scratch, achieves $\textbf{1.72}$ FID with a single function evaluation (1-NFE) on ImageNet 256$\times$256. iMF substantially outperforms prior methods of this kind and closes the gap with multi-step methods while using no distillation. We hope our work will further advance fastforward generative modeling as a stand-alone paradigm.

生成模型快速生成单步采样图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。