arXiv:2510.20771cs.CVcs.LG2025-10被引 39

提出α-Flow优化均值流模型,解决训练冲突问题,提升生成质量。

AlphaFlow: Understanding and Improving MeanFlow Models

  • 统一轨迹流匹配与均值流目标,设计可平滑过渡的课程学习策略
  • 在ImageNet-1K上实现FID 2.58(1-NFE)和2.15(2-NFE)新纪录
  • 适合关注生成模型训练机制与性能优化的研究者

均值流(MeanFlow)近期成为从零训练的少步生成建模的强大框架,但其成功机制尚未完全明晰。本文发现均值流目标天然分解为轨迹流匹配与轨迹一致性两部分,且二者梯度强负相关,导致优化冲突与收敛缓慢。基于此,我们提出α-Flow,一个统一轨迹流匹配、快捷模型与均值流的广义目标族。通过从轨迹流匹配到均值流的平滑课程训练策略,α-Flow解耦冲突目标,实现更快收敛。在使用原始DiT骨干网络从零训练的条件式ImageNet-1K 256x256数据集上,α-Flow在不同规模与设置下持续优于均值流。最大模型α-Flow-XL/2+以原生DiT骨架达成新最优结果,1-NFE时FID为2.58,2-NFE时为2.15。

原文摘要 · Abstract (English)

MeanFlow has recently emerged as a powerful framework for few-step generative modeling trained from scratch, but its success is not yet fully understood. In this work, we show that the MeanFlow objective naturally decomposes into two parts: trajectory flow matching and trajectory consistency. Through gradient analysis, we find that these terms are strongly negatively correlated, causing optimization conflict and slow convergence. Motivated by these insights, we introduce $α$-Flow, a broad family of objectives that unifies trajectory flow matching, Shortcut Model, and MeanFlow under one formulation. By adopting a curriculum strategy that smoothly anneals from trajectory flow matching to MeanFlow, $α$-Flow disentangles the conflicting objectives, and achieves better convergence. When trained from scratch on class-conditional ImageNet-1K 256x256 with vanilla DiT backbones, $α$-Flow consistently outperforms MeanFlow across scales and settings. Our largest $α$-Flow-XL/2+ model achieves new state-of-the-art results using vanilla DiT backbones, with FID scores of 2.58 (1-NFE) and 2.15 (2-NFE).

生成模型流匹配优化改进DiT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。