arXiv:2512.02826cs.LGcs.AI2025-12中稿 · CVPR被引 7

揭示流模型训练的两阶段机制:先导航后精修。

From Navigation to Refinement: Revealing the Two-Stage Nature of Flow-based Diffusion Models through Oracle Velocity

  • 通过解析速度场,发现模型分早期导航与后期精修两阶段
  • 早期跨模式泛化形成整体布局,后期聚焦最近样本细节
  • 解释了时间偏移、无分类器引导等实用技巧的有效性

基于流的扩散模型已成为图像与视频生成的重要范式,但其记忆-泛化行为仍不清晰。本文重新审视流匹配(FM)目标,研究其边缘速度场,该场具有闭式表达,可精确计算理想FM目标。分析表明,流模型天然具备两阶段训练目标:早期由数据模态混合引导,后期受最近数据样本主导。这两个阶段导致不同学习行为:早期为跨模态泛化以构建全局布局,后期则逐步记忆细粒度细节。基于此,我们解释了时间偏移调度、无分类器引导区间及潜在空间设计等实际技术的有效性。本研究深化了对扩散模型训练动态的理解,并为未来架构与算法改进提供了指导原则。

原文摘要 · Abstract (English)

Flow-based diffusion models have emerged as a leading paradigm for training generative models across images and videos. However, their memorization-generalization behavior remains poorly understood. In this work, we revisit the flow matching (FM) objective and study its marginal velocity field, which admits a closed-form expression, allowing exact computation of the oracle FM target. Analyzing this oracle velocity field reveals that flow-based diffusion models inherently formulate a two-stage training target: an early stage guided by a mixture of data modes, and a later stage dominated by the nearest data sample. The two-stage objective leads to distinct learning behaviors: the early navigation stage generalizes across data modes to form global layouts, whereas the later refinement stage increasingly memorizes fine-grained details. Leveraging these insights, we explain the effectiveness of practical techniques such as timestep-shifted schedules, classifier-free guidance intervals, and latent space design choices. Our study deepens the understanding of diffusion model training dynamics and offers principles for guiding future architectural and algorithmic improvements. Our project page is available at: https://maps-research.github.io/from-navigation-to-refinement/.

扩散模型训练机制两阶段

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。