arXiv:2509.10441cs.CV2025-09ICCV被引 4

让图像生成支持任意分辨率,4K出图快过10秒

InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis

  • 用单步生成器解码固定尺寸潜在表示,突破分辨率限制
  • 4K图像生成时间从100秒以上缩短至10秒内
  • 无需重训练模型,适配所有同潜空间扩散模型

任意分辨率图像生成可为不同设备提供一致的视觉体验,在生产与消费场景中应用广泛。当前扩散模型的计算开销随分辨率呈平方级增长,导致4K图像生成延迟超过100秒。为此,我们提出第二代潜空间扩散模型,将扩散模型生成的固定潜表示视为内容表征,并设计一个紧凑的单步生成器,从该固定大小的潜在表示中解码任意分辨率图像。我们提出InfGen,用新生成器替代VAE解码器,实现无需重新训练扩散模型即可生成任意分辨率图像,简化流程、降低计算复杂度,且适用于任何共享相同潜空间的模型。实验表明,InfGen可使多个模型进入任意高分辨率生成时代,同时将4K图像生成时间压缩至10秒以下。

原文摘要 · Abstract (English)

Arbitrary resolution image generation provides a consistent visual experience across devices, having extensive applications for producers and consumers. Current diffusion models increase computational demand quadratically with resolution, causing 4K image generation delays over 100 seconds. To solve this, we explore the second generation upon the latent diffusion models, where the fixed latent generated by diffusion models is regarded as the content representation and we propose to decode arbitrary resolution images with a compact generated latent using a one-step generator. Thus, we present the \textbf{InfGen}, replacing the VAE decoder with the new generator, for generating images at any resolution from a fixed-size latent without retraining the diffusion models, which simplifies the process, reducing computational complexity and can be applied to any model using the same latent space. Experiments show InfGen is capable of improving many models into the arbitrary high-resolution era while cutting 4K image generation time to under 10 seconds.

图像生成扩散模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。