arXiv:2512.18445cs.LG2025-12

解析Transformer为何通用,揭示注意力机制的极限与边界

On the Universality of Transformer Architectures; How Much Attention Is Enough?

  • 从理论角度分析Transformer的表达能力与结构最小化
  • 梳理近年进展,明确哪些结论可靠、哪些易受条件影响
  • 适合对模型原理感兴趣的研宄者和理论学习者

Transformer在大语言模型、计算机视觉和强化学习等多个AI领域至关重要。其广泛应用源于架构所展现的通用性与可扩展性。本文研究Transformer的通用性问题,回顾近期进展,包括结构最小化、近似率等架构优化,并综述当前最先进的成果,以深化理论与实践理解。目标是厘清现有知识中Transformer表达能力的本质,区分稳固保证与脆弱结论,并指明未来理论研究的关键方向。

原文摘要 · Abstract (English)

Transformers are crucial across many AI fields, such as large language models, computer vision, and reinforcement learning. This prominence stems from the architecture's perceived universality and scalability compared to alternatives. This work examines the problem of universality in Transformers, reviews recent progress, including architectural refinements such as structural minimality and approximation rates, and surveys state-of-the-art advances that inform both theoretical and practical understanding. Our aim is to clarify what is currently known about Transformers expressiveness, separate robust guarantees from fragile ones, and identify key directions for future theoretical research.

Transformer理论分析通用性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。