发现自回归视觉生成本质是离散扩散,可提升效率与质量
Scale-Wise VAR is Secretly Discrete Diffusion
- 用马尔可夫注意力掩码揭示自回归生成等价于离散扩散
- 引入扩散优势后,推理速度更快、零样本重建更优
- 适合关注生成模型效率与架构优化的研究者
自回归(AR)Transformer已成为视觉生成的有力范式,因其可扩展性、计算效率及与语言和视觉的统一架构。其中,下一尺度预测的视觉自回归生成(VAR)近期表现出色,甚至超越基于扩散的模型。本文重新审视VAR,发现当使用马尔可夫注意力掩码时,其数学上等价于离散扩散过程。我们将其重新诠释为可扩展的视觉精炼离散扩散(SRDD),建立了自回归Transformer与扩散模型之间的理论桥梁。基于此视角,可直接引入扩散模型的优势,如迭代精炼,并减少架构低效,实现更快收敛、更低推理成本和更好的零样本重建。在多个数据集上,该扩散视角均带来效率与生成质量的持续提升。
原文摘要 · Abstract (English)
Autoregressive (AR) transformers have emerged as a powerful paradigm for visual generation, largely due to their scalability, computational efficiency and unified architecture with language and vision. Among them, next scale prediction Visual Autoregressive Generation (VAR) has recently demonstrated remarkable performance, even surpassing diffusion-based models. In this work, we revisit VAR and uncover a theoretical insight: when equipped with a Markovian attention mask, VAR is mathematically equivalent to a discrete diffusion. We term this reinterpretation as Scalable Visual Refinement with Discrete Diffusion (SRDD), establishing a principled bridge between AR transformers and diffusion models. Leveraging this new perspective, we show how one can directly import the advantages of diffusion such as iterative refinement and reduce architectural inefficiencies into VAR, yielding faster convergence, lower inference cost, and improved zero-shot reconstruction. Across multiple datasets, we show that the diffusion based perspective of VAR leads to consistent gains in efficiency and generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。