arXiv:2506.05240cs.LGcs.CV2025-06被引 2

用流模型当先验,高效对齐潜在空间分布。

Aligning Latent Spaces with Flow Priors

  • 用流模型预训练捕捉目标分布,作为潜在空间正则化先验
  • 无需求解ODE或计算似然,优化更高效,理论可证明等价于最大化对数似然下界
  • 在ImageNet上验证有效,适合需高效对齐潜在空间的研究者

本文提出一种新框架,通过基于流的生成模型作为先验,将可学习的潜在空间对齐到任意目标分布。方法首先在目标特征上预训练流模型以捕获底层分布,固定后的流模型通过一种对齐损失来正则化潜在空间,该损失重构了流匹配目标,将潜在变量视为优化目标。我们严格证明最小化此对齐损失等价于最大化潜在变量在目标分布下的变分对数似然下界,且无需进行耗时的似然评估或优化过程中的ODE求解。在受控设置中,我们验证了对齐损失曲面近似于目标分布的负对数似然。进一步在ImageNet上进行了大规模图像生成实验,采用多种目标分布,并包含详尽的分析与消融研究。理论与实证双重验证表明,该框架为潜在空间对齐开辟了新路径。

原文摘要 · Abstract (English)

This paper presents a novel framework for aligning learnable latent spaces to arbitrary target distributions by leveraging flow-based generative models as priors. Our method first pretrains a flow model on the target features to capture the underlying distribution. This fixed flow model subsequently regularizes the latent space via an alignment loss, which reformulates the flow matching objective to treat the latents as optimization targets. We formally prove that minimizing this alignment loss establishes a computationally tractable surrogate objective for maximizing a variational lower bound on the log-likelihood of latents under the target distribution. Notably, the proposed method eliminates computationally expensive likelihood evaluations and avoids ODE solving during optimization. As a proof of concept, we demonstrate in a controlled setting that the alignment loss landscape closely approximates the negative log-likelihood of the target distribution. We further validate the effectiveness of our approach through large-scale image generation experiments on ImageNet with diverse target distributions, accompanied by detailed discussions and ablation studies. With both theoretical and empirical validation, our framework paves a new way for latent space alignment.

潜在空间对齐流模型生成模型高效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。