arXiv:2511.19065cs.CVcs.AI2025-11被引 2

改进均流训练,让生成模型更快更准

Understanding, Accelerating, and Improving MeanFlow Training

  • 先学瞬时速度,再逐步学习长间隔平均速度
  • 新方法使图像生成FID降至2.87,优于原版3.43
  • 适合追求高效生成的扩散模型研究者

MeanFlow通过联合学习瞬时与平均速度场,在少步内实现高质量生成,但其训练机制尚不清晰。我们分析发现:(i) 学习平均速度需以良好瞬时速度为基础;(ii) 小时间间隔下平均速度能促进瞬时速度学习,但间隔增大则导致性能下降;(iii) 长间隔平均速度的平滑学习依赖于准确的瞬时速度和小间隔平均速度的前期形成。基于此,我们设计新训练策略:优先加速瞬时速度形成,再转向长间隔平均速度。该方法显著提升收敛速度与生成质量:相同DiT-XL架构下,1步生成在ImageNet 256x256上FID达2.87(原版为3.43);或以2.5倍更短训练时间达到原版性能,或使用更小的DiT-L模型达成同等效果。

原文摘要 · Abstract (English)

MeanFlow promises high-quality generative modeling in few steps, by jointly learning instantaneous and average velocity fields. Yet, the underlying training dynamics remain unclear. We analyze the interaction between the two velocities and find: (i) well-established instantaneous velocity is a prerequisite for learning average velocity; (ii) learning of instantaneous velocity benefits from average velocity when the temporal gap is small, but degrades as the gap increases; and (iii) task-affinity analysis indicates that smooth learning of large-gap average velocities, essential for one-step generation, depends on the prior formation of accurate instantaneous and small-gap average velocities. Guided by these observations, we design an effective training scheme that accelerates the formation of instantaneous velocity, then shifts emphasis from short- to long-interval average velocity. Our enhanced MeanFlow training yields faster convergence and significantly better few-step generation: With the same DiT-XL backbone, our method reaches an impressive FID of 2.87 on 1-NFE ImageNet 256x256, compared to 3.43 for the conventional MeanFlow baseline. Alternatively, our method matches the performance of the MeanFlow baseline with 2.5x shorter training time, or with a smaller DiT-L backbone.

扩散模型生成模型训练加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。