arXiv:2502.06167cs.LGcs.AI2025-02被引 13

证明了简单视觉自回归变压器能逼近任意图像映射,为高效图像生成提供理论基础。

Universal Approximation of Visual Autoregressive Transformers

  • 用单头自注意力和单层插值构建简单模型,实现图像到图像的通用逼近。
  • 理论证明该模型可逼近任意满足利普希茨条件的图像映射函数。
  • 适用于图像生成、高效设计等场景,为后续模型优化指明方向。

我们研究基于Transformer的基座模型的基本极限,将分析扩展至视觉自回归(VAR)Transformer。VAR代表了利用新颖、可扩展的粗到细「下一尺度预测」框架生成图像的重要进展。这些模型设定了新的质量标准,在图像合成任务中超越所有先前方法(包括扩散Transformer),并达到当前最优性能。主要贡献在于证明:对于具有单头自注意力层和单插值层的单头VAR Transformer,其具备通用逼近能力。从统计角度看,我们证明此类简单模型是任意图像到图像的利普希茨函数的通用逼近器。此外,我们还展示了基于流的自回归Transformer也具备类似的逼近能力。这些结果为有效且计算高效的VAR Transformer策略提供了重要设计原则,可用于扩展其在图像生成及其他相关领域的应用。

原文摘要 · Abstract (English)

We investigate the fundamental limits of transformer-based foundation models, extending our analysis to include Visual Autoregressive (VAR) transformers. VAR represents a big step toward generating images using a novel, scalable, coarse-to-fine ``next-scale prediction'' framework. These models set a new quality bar, outperforming all previous methods, including Diffusion Transformers, while having state-of-the-art performance for image synthesis tasks. Our primary contributions establish that, for single-head VAR transformers with a single self-attention layer and single interpolation layer, the VAR Transformer is universal. From the statistical perspective, we prove that such simple VAR transformers are universal approximators for any image-to-image Lipschitz functions. Furthermore, we demonstrate that flow-based autoregressive transformers inherit similar approximation capabilities. Our results provide important design principles for effective and computationally efficient VAR Transformer strategies that can be used to extend their utility to more sophisticated VAR models in image generation and other related areas.

图像生成视觉变压器理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。