arXiv:2412.07972cs.LGstat.ML2024-12

提出时间调度方法,让生成模型分阶段学习概率与方差。

Phase-aware Training Schedule Simplifies Learning in Flow-Based Generative Models

  • 用时间拉伸解决高维下相位消失问题,使训练分两阶段进行。
  • 发现模型在不同阶段只学习对应参数,自动简化结构。
  • 针对真实数据设计高效训练策略,提升特征学习效率。

我们分析了用于参数化流模型以从高维高斯混合分布采样的双层自编码器的训练过程。先前研究指出,当维度趋于无穷时,模式间相对概率的学习相位会消失,除非采用合适的时间调度。我们引入时间拉伸机制解决了此问题。这使得我们能够刻画学习到的速度场:第一阶段学习各模式的概率,第二阶段学习各模式的方差。我们发现,表示速度场的自编码器会通过仅估计每个阶段相关参数来自我简化。转向真实数据,我们提出一种方法,对给定特征识别出训练精度提升最显著的时间区间。由于实践者通常均匀采样训练时间,我们的方法实现了更高效的训练。初步实验验证了该方法的有效性。

原文摘要 · Abstract (English)

We analyze the training of a two-layer autoencoder used to parameterize a flow-based generative model for sampling from a high-dimensional Gaussian mixture. Previous work shows that the phase where the relative probability between the modes is learned disappears as the dimension goes to infinity without an appropriate time schedule. We introduce a time dilation that solves this problem. This enables us to characterize the learned velocity field, finding a first phase where the probability of each mode is learned and a second phase where the variance of each mode is learned. We find that the autoencoder representing the velocity field learns to simplify by estimating only the parameters relevant to each phase. Turning to real data, we propose a method that, for a given feature, finds intervals of time where training improves accuracy the most on that feature. Since practitioners take a uniform distribution over training times, our method enables more efficient training. We provide preliminary experiments validating this approach.

流模型训练调度生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。