arXiv:2502.09611cs.LGcs.CV2025-02被引 9

为流模型设计条件先验,加速生成并提升质量

Designing a Conditional Prior Distribution for Flow-Based Generative Models

  • 根据输入条件生成数据空间中的平均点,构建非平凡先验
  • 采样步数更少,FID、KID和CLIP得分显著优于基线
  • 适合追求高效高质量生成的开发者与研究者

流式生成模型在文本到图像等条件生成任务中表现优异,但现有方法将通用单峰噪声分布映射至目标数据分布的特定模式,导致初始分布中每一点可映射到目标分布中任意点,产生较长平均路径。为此,本文利用条件流模型的一个未被充分利用特性:可设计非平凡先验分布。给定输入条件(如文本提示),首先将其映射到数据空间中一个代表“平均”数据点,该点与同条件模式下的所有数据点的平均距离最小。随后,采用流匹配公式,将围绕该点的参数化分布样本映射至条件目标分布。实验表明,相比基线方法,本方法显著提升训练速度与生成效率(FID、KID及CLIP对齐分数),以更少采样步数生成高质量样本。

原文摘要 · Abstract (English)

Flow-based generative models have recently shown impressive performance for conditional generation tasks, such as text-to-image generation. However, current methods transform a general unimodal noise distribution to a specific mode of the target data distribution. As such, every point in the initial source distribution can be mapped to every point in the target distribution, resulting in long average paths. To this end, in this work, we tap into a non-utilized property of conditional flow-based models: the ability to design a non-trivial prior distribution. Given an input condition, such as a text prompt, we first map it to a point lying in data space, representing an ``average" data point with the minimal average distance to all data points of the same conditional mode (e.g., class). We then utilize the flow matching formulation to map samples from a parametric distribution centered around this point to the conditional target distribution. Experimentally, our method significantly improves training times and generation efficiency (FID, KID and CLIP alignment scores) compared to baselines, producing high quality samples using fewer sampling steps.

流模型条件生成高效采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。