arXiv:2511.16117cs.CV2025-11被引 1

让图像生成不再受分辨率限制,灵活调整质量与计算量。

Decoupling Complexity from Scale in Latent Diffusion Model

  • 构建分层无尺度的潜在空间,用多级令牌表示内容复杂度。
  • 固定潜在表示下支持任意分辨率和帧率生成,性能媲美顶尖模型。
  • 支持从粗到细渐进生成,适合需要灵活控制画质的场景。

现有潜空间扩散模型通常将规模与内容复杂度耦合,用更多潜在标记表示高分辨率图像或高帧率视频。然而,表示视觉数据所需的潜在容量主要取决于内容复杂度,规模仅作为上限。基于此观察,我们提出DCS-LDM,一种将信息复杂度与规模解耦的新型视觉生成范式。DCS-LDM构建分层、与尺度无关的潜在空间,通过多级标记建模样本复杂度,并在固定潜在表示下支持任意分辨率和帧率的解码。该潜在空间使DCS-LDM实现灵活的计算-质量权衡。此外,通过跨层级分解结构与细节信息,支持渐进式粗到精生成。实验表明,DCS-LDM在多种尺度和视觉质量下表现接近当前最优方法。

原文摘要 · Abstract (English)

Existing latent diffusion models typically couple scale with content complexity, using more latent tokens to represent higher-resolution images or higher-frame rate videos. However, the latent capacity required to represent visual data primarily depends on content complexity, with scale serving only as an upper bound. Motivated by this observation, we propose DCS-LDM, a novel paradigm for visual generation that decouples information complexity from scale. DCS-LDM constructs a hierarchical, scale-independent latent space that models sample complexity through multi-level tokens and supports decoding to arbitrary resolutions and frame rates within a fixed latent representation. This latent space enables DCS-LDM to achieve a flexible computation-quality tradeoff. Furthermore, by decomposing structural and detailed information across levels, DCS-LDM supports a progressive coarse-to-fine generation paradigm. Experimental results show that DCS-LDM delivers performance comparable to state-of-the-art methods while offering flexible generation across diverse scales and visual qualities.

扩散模型图像生成潜空间可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。