arXiv:2605.07915cs.CV2026-05被引 5

提出新型自编码器,让潜在空间更利于生成高质量图像。

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion

论文配图:What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion
图 1 · 摘自论文原文
  • 通过显式优化空间结构、局部连续性和全局语义构建生成友好潜空间。
  • 在ImageNet上实现1.03的gFID新纪录,训练速度提升13倍。
  • 适合追求高效高质图像生成的研究者与开发者使用。

潜空间中的分词器是潜扩散模型的关键组件,决定了扩散模型运作的潜在空间。然而,现有分词器主要关注重建保真度或继承预训练表征,未明确何种潜空间真正有利于生成建模。本文从潜流形组织角度研究该问题,通过构建受控的分词器变体,识别出三个生成友好潜流形的关键属性:一致的空间结构、局部流形连续性与全局流形语义。研究发现这些属性比重建保真度更与下游生成质量相关。基于此,提出先验对齐自编码器(PAE),显式塑造潜流形而非依赖重建或继承间接产生。PAE利用来自变分流形模型(VFMs)的精炼先验和基于扰动的正则化,将空间结构、局部连续性与全局语义转化为显式训练目标。在ImageNet 256x256上,PAE相较现有分词器提升训练效率与生成质量,达到与RAE相当的性能,但收敛速度提升高达13倍,并实现新的gFID记录1.03。结果凸显了潜流形组织对潜扩散模型的重要性。

原文摘要 · Abstract (English)

Tokenizers are a crucial component of latent diffusion models, as they define the latent space in which diffusion models operate. However, existing tokenizers are primarily designed to improve reconstruction fidelity or inherit pretrained representations, leaving unclear what kind of latent space is truly friendly for generative modeling. In this paper, we study this question from the perspective of latent manifold organization. By constructing controlled tokenizer variants, we identify three key properties of a diffusion-friendly latent manifold: coherent spatial structure, local manifold continuity, and global manifold semantics. We find that these properties are more consistent with downstream generation quality than reconstruction fidelity. Motivated by this finding, we propose the Prior-Aligned AutoEncoder (PAE), which explicitly shapes the latent manifold instead of leaving diffusion-friendly manifold to emerge indirectly from reconstruction or inheritance. Specifically, PAE leverages refined priors derived from VFMs and perturbation-based regularization to turn spatial structure, local continuity, and global semantics into explicit training objectives. On ImageNet 256x256, PAE improves both training efficiency and generation quality over existing tokenizers, reaching performance comparable to RAE with up to 13x faster convergence under the same training setup and achieving a new state-of-the-art gFID of 1.03. These results highlight the importance of organizing the latent manifold for latent diffusion models.

潜空间扩散模型自编码器图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。