arXiv:2511.13387cs.CVcs.AI2025-11

用预训练扩散模型高效生成图像令牌,兼容主流生成架构。

Generalized Denoising Diffusion Codebook Models (gDDCM): Tokenizing images using a pre-trained diffusion model

  • 提出统一框架与反向追踪采样策略,适配连续时间扩散模型。
  • 在CIFAR10和LSUN Bedroom上重建质量显著优于传统DDCM。
  • 支持流匹配与一致性模型,提升高噪声区域采样效率。

去噪扩散模型已成为图像生成的主流范式。将图像数据离散化为令牌是有效整合图像与Transformer等架构的关键步骤。尽管去噪扩散码本模型(DDCM)开创性地利用预训练扩散模型进行图像令牌化,但其严格依赖传统的离散时间DDPM架构,无法适配现代连续时间变体(如流匹配和一致性模型),且在高噪声区域存在采样效率低下问题。为此,本文提出广义去噪扩散码本模型(gDDCM)。我们建立了统一的理论框架,并引入通用的“去噪并回溯”采样策略。通过结合确定性常微分方程去噪步骤与残差对齐噪声注入步骤,解决了适应性难题。此外,引入回溯参数 $p$,显著提升了令牌化能力。在CIFAR10和LSUN Bedroom数据集上的大量实验表明,gDDCM实现了与主流扩散模型的全面兼容,在重建质量和感知保真度方面显著优于DDCM。

原文摘要 · Abstract (English)

Denoising diffusion models have emerged as a dominant paradigm in image generation. Discretizing image data into tokens is a critical step for effectively integrating images with Transformer and other architectures. Although the Denoising Diffusion Codebook Models (DDCM) pioneered the use of pre-trained diffusion models for image tokenization, it strictly relies on the traditional discrete-time DDPM architecture. Consequently, it fails to adapt to modern continuous-time variants-such as Flow Matching and Consistency Models-and suffers from inefficient sampling in high-noise regions. To address these limitations, this paper proposes the Generalized Denoising Diffusion Codebook Models (gDDCM). We establish a unified theoretical framework and introduce a generic "De-noise and Back-trace" sampling strategy. By integrating a deterministic ODE denoising step with a residual-aligned noise injection step, our method resolves the challenge of adaptation. Furthermore, we introduce a backtracking parameter $p$ and significantly enhance tokenization ability. Extensive experiments on CIFAR10 and LSUN Bedroom datasets demonstrate that gDDCM achieves comprehensive compatibility with mainstream diffusion variants and significantly outperforms DDCM in terms of reconstruction quality and perceptual fidelity.

扩散模型图像令牌化生成模型连续时间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。