首次揭示视觉自回归模型的表达能力上限,虽效果好但受限于简单电路。
Circuit Complexity Bounds for Visual Autoregressive Model
- 用电路复杂度分析方法研究视觉自回归模型
- 证明其可被多项式精度的TC⁰电路模拟,隐藏维度不超过O(n)
- 为高效模型设计提供理论依据,适合关注生成模型原理的研究者
理解特定模型的表达能力对于把握其性能局限至关重要。近期多项研究已建立Transformer架构的电路复杂度界。此外,视觉自回归(VAR)模型在图像生成领域崭露头角,生成高质量图像的能力超越了此前的扩散Transformer等技术。本文研究了VAR模型的电路复杂度,并建立了相应界限。主要结果表明,VAR模型等价于一个均匀的$\mathsf{TC}^0$阈值电路,其隐藏维度$d \leq O(n)$,且具有$\mathrm{poly}(n)$精度。这是首个严格揭示尽管表现优异,但VAR模型表达能力存在内在限制的研究。我们相信这些发现将为理解此类模型的本质约束提供重要洞见,并指导未来更高效、更具表达力架构的发展。
原文摘要 · Abstract (English)
Understanding the expressive ability of a specific model is essential for grasping its capacity limitations. Recently, several studies have established circuit complexity bounds for Transformer architecture. Besides, the Visual AutoRegressive (VAR) model has risen to be a prominent method in the field of image generation, outperforming previous techniques, such as Diffusion Transformers, in generating high-quality images. We investigate the circuit complexity of the VAR model and establish a bound in this study. Our primary result demonstrates that the VAR model is equivalent to a simulation by a uniform $\mathsf{TC}^0$ threshold circuit with hidden dimension $d \leq O(n)$ and $\mathrm{poly}(n)$ precision. This is the first study to rigorously highlight the limitations in the expressive power of VAR models despite their impressive performance. We believe our findings will offer valuable insights into the inherent constraints of these models and guide the development of more efficient and expressive architectures in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。