arXiv:2608.12671cs.AIcs.CC2026-08

用电路复杂度分析Transformer的表达能力,揭示其语言建模极限。

On the Expressive Power of Transformers

论文配图:On the Expressive Power of Transformers
图 1 · 摘自论文原文
  • 通过电路复杂度框架,对比Transformer与经典计算模型
  • 证明Transformer可模拟特定类别的有界深度电路
  • 为大模型表达能力提供理论边界,适合理论研究者

多层Transformer是当前几乎所有大型语言模型的核心组件。由于其广泛应用和强大计算能力,越来越多的研究试图通过与理论计算机科学中长期研究的标准计算模型对比,精确刻画Transformer作为语言识别器的表达能力。在此过程中,电路复杂度已成为分析Transformer表达能力的“合适”分支——因为将Transformer按注意力、精度等资源参数化,可直接与按门类型、规模、深度等资源参数化的电路类别进行比较。本文综述了利用电路复杂度概念与方法,刻画Transformer表达能力的若干关键结果。

原文摘要 · Abstract (English)

Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today. Because of their ubiquity and computational capability, there is a rapidly growing body of work that aims to precisely calibrate the expressive power of transformers as language recognizers by comparing them against standard models of computation studied for decades by the theoretical computer science community. In this endeavor, circuit complexity has by and large emerged as the "correct" branch of computational complexity to analyze the expressive power of transformers; the reason is that parameterizing transformers by the various resources they use, such as attention and precision, leads to direct comparisons with different classes of circuits parameterized by resources such as type of gates, size, and depth. Here, we present an overview of selected results that delineate the expressive power of transformers using concepts and methods from circuit complexity.

Transformer表达能力理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。