让生成模型的潜在空间更懂语义,生成效果提升15%。
Exploring Representation-Aligned Latent Space for Better Generation
- 用语义先验构建对齐表征的潜在空间,替代传统压缩式编码。
- 在该空间训练的DiT/SiT模型FID降低15%,生成质量显著提升。
- 适合关注图像生成与下游感知任务优化的研究者。
生成模型是建模真实世界的重要工具,主流扩散模型(尤其是基于潜在扩散模型范式)已在图像、视频合成等任务中取得显著进展。这类模型通常通过变分自编码器(VAE)训练,与潜在表示而非原始样本交互。尽管该范式加速了训练与推理,但生成质量受限于潜在表示的质量。传统VAE潜在变量常被视为像素空间中的空间压缩,缺乏显式的语义表征,而语义信息对建模真实世界至关重要。本文提出ReaLS(Representation-Aligned Latent Space),通过融入语义先验以提升生成性能。大量实验表明,基于ReaLS训练的基础DiT与SiT模型在FID指标上实现15%的提升。此外,增强后的语义潜在空间还显著提升了分割、深度估计等感知下游任务的表现。
原文摘要 · Abstract (English)
Generative models serve as powerful tools for modeling the real world, with mainstream diffusion models, particularly those based on the latent diffusion model paradigm, achieving remarkable progress across various tasks, such as image and video synthesis. Latent diffusion models are typically trained using Variational Autoencoders (VAEs), interacting with VAE latents rather than the real samples. While this generative paradigm speeds up training and inference, the quality of the generated outputs is limited by the latents' quality. Traditional VAE latents are often seen as spatial compression in pixel space and lack explicit semantic representations, which are essential for modeling the real world. In this paper, we introduce ReaLS (Representation-Aligned Latent Space), which integrates semantic priors to improve generation performance. Extensive experiments show that fundamental DiT and SiT trained on ReaLS can achieve a 15% improvement in FID metric. Furthermore, the enhanced semantic latent space enables more perceptual downstream tasks, such as segmentation and depth estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。