揭示Transformer在符号推理中的三大理论瓶颈,解释为何难做精确计算。
Barriers to Discrete Reasoning with Transformers: A Survey Across Depth, Exactness, and Bandwidth
- 从电路、逼近和通信三理论视角分析Transformer的结构缺陷
- 深度受限、难以处理不连续函数、跨标记通信效率低是主要障碍
- 适合研究大模型原理或设计新架构的学者参考
Transformer已成为自然语言处理、视觉等序列建模任务的核心架构,但在算术、逻辑推理和算法组合等离散推理任务中仍存在理论局限性。本文综合电路复杂性、逼近理论和通信复杂性三个理论视角,系统阐述Transformer在执行符号计算时面临的结构性与计算性障碍。通过整合已有成果,揭示当前架构即使在模式匹配和插值上表现优异,仍难以实现精确的离散算法。文中回顾关键定义、经典结论与典型案例,指出深度约束、不连续性逼近困难以及标记间通信瓶颈等问题。最后讨论对模型设计的启示,并提出突破这些基础限制的潜在方向。
原文摘要 · Abstract (English)
Transformers have become the foundational architecture for a broad spectrum of sequence modeling applications, underpinning state-of-the-art systems in natural language processing, vision, and beyond. However, their theoretical limitations in discrete reasoning tasks, such as arithmetic, logical inference, and algorithmic composition, remain a critical open problem. In this survey, we synthesize recent studies from three theoretical perspectives: circuit complexity, approximation theory, and communication complexity, to clarify the structural and computational barriers that transformers face when performing symbolic computations. By connecting these established theoretical frameworks, we provide an accessible and unified account of why current transformer architectures struggle to implement exact discrete algorithms, even as they excel at pattern matching and interpolation. We review key definitions, seminal results, and illustrative examples, highlighting challenges such as depth constraints, difficulty approximating discontinuities, and bottlenecks in inter-token communication. Finally, we discuss implications for model design and suggest promising directions for overcoming these foundational limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。