提出一种高效稳定的单步生成模型,提升生成速度与泛化能力。
Modular MeanFlow: Towards Stable and Scalable One-Step Generative Modeling
- 基于平均速度场学习,通过微分恒等式设计损失函数。
- 在图像与轨迹生成任务中实现高质量样本与稳定收敛。
- 适合数据少或分布外场景,无需高阶导数计算。
单步生成建模旨在一次函数评估中生成高质量数据,显著优于传统扩散或流模型。本文提出模块化平均流(Modular MeanFlow, MMF),一种灵活且理论严谨的时均速度场学习方法。该方法基于瞬时与平均速度间的微分恒等式,推导出一类损失函数,并引入梯度调制机制,实现稳定训练而不损失表达能力。进一步提出课程式预热策略,平滑过渡至全可微训练。MMF统一并推广了现有基于一致性与流匹配的方法,同时避免昂贵的高阶导数计算。在图像合成与轨迹建模任务上的实验证明,MMF在低数据或分布外设置下仍具备竞争力的样本质量、鲁棒收敛性与强泛化能力。
原文摘要 · Abstract (English)
One-step generative modeling seeks to generate high-quality data samples in a single function evaluation, significantly improving efficiency over traditional diffusion or flow-based models. In this work, we introduce Modular MeanFlow (MMF), a flexible and theoretically grounded approach for learning time-averaged velocity fields. Our method derives a family of loss functions based on a differential identity linking instantaneous and average velocities, and incorporates a gradient modulation mechanism that enables stable training without sacrificing expressiveness. We further propose a curriculum-style warmup schedule to smoothly transition from coarse supervision to fully differentiable training. The MMF formulation unifies and generalizes existing consistency-based and flow-matching methods, while avoiding expensive higher-order derivatives. Empirical results across image synthesis and trajectory modeling tasks demonstrate that MMF achieves competitive sample quality, robust convergence, and strong generalization, particularly under low-data or out-of-distribution settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。