arXiv:2602.17270cs.LGcs.CV2026-02被引 16

统一潜在空间让生成模型更高效,图像视频质量双提升。

Unified Latents (UL): How to train your latents

  • 用扩散先验约束编码器输出噪声,简化训练目标。
  • 图像上FID达1.4,视频上FVD低至1.3,性能领先。
  • 训练算力更低,适合追求高效生成的开发者。

我们提出统一潜在空间(UL),一种联合利用扩散先验正则化潜在表示,并由扩散模型解码的框架。通过将编码器输出噪声与先验最小噪声水平关联,获得一个简单训练目标,该目标对潜在比特率提供紧致上界。在ImageNet-512上,方法实现1.4的竞争力FID,同时保持高重建质量(PSNR),且训练所需浮点运算量低于基于Stable Diffusion潜在空间训练的模型。在Kinetics-600上,将FVD新纪录降至1.3。

原文摘要 · Abstract (English)

We present Unified Latents (UL), a framework for learning latent representations that are jointly regularized by a diffusion prior and decoded by a diffusion model. By linking the encoder's output noise to the prior's minimum noise level, we obtain a simple training objective that provides a tight upper bound on the latent bitrate. On ImageNet-512, our approach achieves competitive FID of 1.4, with high reconstruction quality (PSNR) while requiring fewer training FLOPs than models trained on Stable Diffusion latents. On Kinetics-600, we set a new state-of-the-art FVD of 1.3.

潜在空间扩散模型高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。