arXiv:2603.10365cs.CV2026-03被引 2

用几何自编码器提升扩散模型的生成质量与效率

Geometric Autoencoder for Diffusion Models

  • 基于视觉基础模型构建优化的低维语义监督目标
  • 在ImageNet-1K上80轮训练即达gFID 1.82,超越现有方法
  • 适合追求高效高质图像生成的研究者与开发者

潜在扩散模型已在高分辨率视觉生成中达到新基准。引入视觉基础模型先验可提升生成效率,但现有潜在设计仍多为经验性。这些方法常难以兼顾语义区分性、重建保真度与潜在紧凑性。本文提出几何自编码器(GAE),系统性解决上述挑战。通过分析多种对齐范式,GAE从视觉基础模型中构建优化的低维语义监督目标,指导自编码器学习。此外,采用潜在归一化替代标准VAE中的严格KL散度,使潜在流形更稳定且专为扩散学习优化。为保障高强度噪声下的鲁棒重建,GAE引入动态噪声采样机制。实验表明,GAE在ImageNet-1K $256 \times 256$ 基准上表现卓越,不使用无分类器引导时,80轮训练即达gFID 1.82,800轮达1.31,显著优于现有最先进方法。除生成质量外,GAE还实现了压缩率、语义深度与重建稳定性之间的优异平衡。结果验证了设计合理性,为潜在扩散建模提供了新范式。代码与模型已公开于https://github.com/sii-research/GAE。

原文摘要 · Abstract (English)

Latent diffusion models have established a new state-of-the-art in high-resolution visual generation. Integrating Vision Foundation Model priors improves generative efficiency, yet existing latent designs remain largely heuristic. These approaches often struggle to unify semantic discriminability, reconstruction fidelity, and latent compactness. In this paper, we propose Geometric Autoencoder (GAE), a principled framework that systematically addresses these challenges. By analyzing various alignment paradigms, GAE constructs an optimized low-dimensional semantic supervision target from VFMs to provide guidance for the autoencoder. Furthermore, we leverage latent normalization that replaces the restrictive KL-divergence of standard VAEs, enabling a more stable latent manifold specifically optimized for diffusion learning. To ensure robust reconstruction under high-intensity noise, GAE incorporates a dynamic noise sampling mechanism. Empirically, GAE achieves compelling performance on the ImageNet-1K $256 \times 256$ benchmark, reaching a gFID of 1.82 at only 80 epochs and 1.31 at 800 epochs without Classifier-Free Guidance, significantly surpassing existing state-of-the-art methods. Beyond generative quality, GAE establishes a superior equilibrium between compression, semantic depth and robust reconstruction stability. These results validate our design considerations, offering a promising paradigm for latent diffusion modeling. Code and models are publicly available at https://github.com/sii-research/GAE.

扩散模型自编码器图像生成潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。