arXiv:2602.05951cs.CVcs.AI2026-02被引 5

优化条件依赖的源分布,让流匹配生成更高效稳定。

Better Source, Better Flow: Learning Condition-Dependent Source Distribution for Flow Matching

  • 学习与条件相关的源分布,更好利用文本等输入信号。
  • 实验显示FID指标收敛速度最快提升3倍,生成质量显著提高。
  • 适合关注生成效率与稳定性的扩散模型研究者参考。

流匹配近年来成为文本到图像生成的有前景替代方案。尽管其可灵活选择任意源分布,但现有方法多沿用扩散模型中的标准高斯分布,极少将源分布本身作为优化目标。本文表明,在现代文本到图像系统中,合理设计源分布不仅可行且有益。我们提出在流匹配目标下学习条件依赖的源分布,以更好利用丰富的条件信号。识别出直接将条件引入源分布会引发分布坍塌和不稳定性等关键问题,并证明适当的方差正则化与源-目标方向对齐对稳定有效学习至关重要。进一步分析了目标表示空间的选择如何影响结构化源的流匹配性能,揭示了此类设计最有效的场景。在多个文本到图像基准上广泛实验表明,性能持续且稳健提升,包括FID收敛速度最快提升3倍,凸显了条件流匹配中合理源分布设计的实际优势。

原文摘要 · Abstract (English)

Flow matching has recently emerged as a promising alternative to diffusion-based generative models, particularly for text-to-image generation. Despite its flexibility in allowing arbitrary source distributions, most existing approaches rely on a standard Gaussian distribution, a choice inherited from diffusion models, and rarely consider the source distribution itself as an optimization target in such settings. In this work, we show that principled design of the source distribution is not only feasible but also beneficial at the scale of modern text-to-image systems. Specifically, we propose learning a condition-dependent source distribution under flow matching objective that better exploit rich conditioning signals. We identify key failure modes that arise when directly incorporating conditioning into the source, including distributional collapse and instability, and show that appropriate variance regularization and directional alignment between source and target are critical for stable and effective learning. We further analyze how the choice of target representation space impacts flow matching with structured sources, revealing regimes in which such designs are most effective. Extensive experiments across multiple text-to-image benchmarks demonstrate consistent and robust improvements, including up to a 3x faster convergence in FID, highlighting the practical benefits of a principled source distribution design for conditional flow matching.

流匹配文本生成生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。