arXiv:2510.05364cs.CL2025-10被引 6

挑战注意力瓶颈,探索更快的序列模型替代方案

The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures

  • 对比分析多种非二次复杂度的序列建模架构
  • 揭示纯注意力模型在长序列下的计算瓶颈
  • 适合关注模型效率与架构演进的研究者

Transformer 在过去七年中主导了序列处理任务,尤其是在语言建模方面。然而,其注意力机制固有的二次复杂度随着上下文长度增加成为显著瓶颈。本文综述了近期克服这一瓶颈的努力,包括(次二次)注意力变体、循环神经网络、状态空间模型以及混合架构的进展。我们从计算与内存复杂度、基准测试结果及根本限制等多个角度,批判性分析这些方法,评估纯注意力 Transformer 的主导地位是否即将受到挑战。

原文摘要 · Abstract (English)

Transformers have dominated sequence processing tasks for the past seven years -- most notably language modeling. However, the inherent quadratic complexity of their attention mechanism remains a significant bottleneck as context length increases. This paper surveys recent efforts to overcome this bottleneck, including advances in (sub-quadratic) attention variants, recurrent neural networks, state space models, and hybrid architectures. We critically analyze these approaches in terms of compute and memory complexity, benchmark results, and fundamental limitations to assess whether the dominance of pure-attention transformers may soon be challenged.

Transformer序列建模注意力机制架构演进

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。