arXiv:2410.02548stat.MLcs.LG2024-10被引 7

分步学习生成流模型,提升训练效率与生成质量

Local Flow Matching Generative Models

  • 将生成过程分解为多步局部流匹配,逐步逼近数据到噪声的路径
  • 在表格、图像和机器人策略生成任务上性能媲美传统流匹配方法
  • 支持模型蒸馏,可加速生成且保持高质量结果

流匹配(Flow Matching, FM)是一种无需模拟的连续可逆流学习方法,用于从噪声生成数据。受扩散过程作为梯度流的变分性质启发,本文提出分步流匹配(Local Flow Matching, LFM),通过一系列子模型逐步学习,每个子模型仅匹配至时间步长的扩散过程。由于每一步中待插值的分布更接近,模型规模更小,训练更高效。该变分视角还使得我们能够基于$χ^2$-散度,证明所提流模型的生成保证,利用扩散过程的压缩性。实践中,LFM的分步结构天然适合模型蒸馏,多种蒸馏技术可应用于加速生成。实验表明,LFM在无条件生成表格与图像数据集,以及条件生成机器人操作策略任务上,表现与传统流匹配相当。

原文摘要 · Abstract (English)

Flow Matching (FM) is a simulation-free method for learning a continuous, invertible flow that interpolates between two distributions, and in particular generates data from noise. Inspired by the variational nature of the diffusion process as a gradient flow, we introduce a stepwise FM model, Local Flow Matching (LFM), which sequentially learns a sequence of FM submodels, each matching a diffusion process up to the time-step size in the data-to-noise direction. In each step, the two distributions to be interpolated by the sub-flow model are closer than those in the full-flow matching model, which interpolates data to noise distributions, enabling smaller models with more efficient training. This variational perspective also allows us to prove a theoretical generation guarantee for the proposed flow model in terms of the $χ^2$-divergence between the generated and true data distributions, leveraging the contraction property of the diffusion process. In practice, the stepwise structure of LFM is naturally amenable to model distillation, and various distillation techniques can be applied to accelerate generation. We empirically demonstrate that LFM achieves competitive generative performance compared to FM on unconditional generation of tabular and image datasets, and on conditional generation of robotic manipulation policies.

流匹配生成模型扩散模型模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。