将视觉自回归模型重新理解为分层迭代重构,揭示其高效高质量的本质。
Multi-scale Autoregressive Models are Laplacian, Discrete, and Latent Diffusion Models in Disguise
- 将自回归建模视为分层潜空间的确定性前向过程与学习的逆过程
- 在少量粗到细步骤内完成生成,兼具效率与高保真度
- 适用于图结构生成与天气预测,可对接扩散模型但保持并行生成
我们重新诠释视觉自回归(VAR)模型为迭代细化模型,以识别驱动其质量-效率权衡的设计选择。不同于仅将其视为逐尺度自回归,我们形式化地将其定义为构建拉普拉斯风格潜空间金字塔的确定性前向过程,以及在少量粗到细步骤中重建样本的已学习反向过程。这一表述使去噪扩散模型之间的联系变得清晰,并突显出三个可能决定VAR效率和样本质量的关键建模选择:在学习的潜空间中进行细化、对代码索引的离散预测、按空间频率分解。我们通过受控实验隔离了每个因素对质量和速度的贡献。此外,我们还讨论了该框架如何适应排列无关的图生成和概率性中距离天气预报,同时在保持少步、尺度并行生成的前提下,为扩散方法提供了实际接触点。
原文摘要 · Abstract (English)
We reinterpret Visual Autoregressive (VAR) models as iterative refinement models to identify which design choices drive their quality-efficiency trade-off. Instead of treating VAR only as next-scale autoregression, we formalise it as a deterministic forward process that builds a Laplacian-style latent pyramid, together with a learned backward process that reconstructs samples in a small number of coarse-to-fine steps. This formulation makes the link to denoising diffusion explicit and highlights three modelling choices that may underlie VAR's efficiency and sample quality: refinement in a learned latent space, discrete prediction over code indices, and decomposition by spatial frequency. We support this view with controlled experiments that isolate the contribution of each factor to quality and speed. We also discuss how the same framework can be adapted to permutation-invariant graph generation and probabilistic medium-range weather forecasting, and how it provides practical points of contact with diffusion methods while preserving few-step, scale-parallel generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。