系统对比21种注意力机制,发现其性能排名不稳,难分高下。
Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off
- 构建EEI评分框架,从效率、表达力、可解释性三方面量化评估
- 蒙特卡洛分析显示超67%样本排名变动超1位,表明细粒度排名不可靠
- 适合关注注意力机制整体演进与研究方向的读者
注意力机制推动机器学习发展已逾十年,从神经机器翻译到具备通用推理能力的语言模型。本综述涵盖四个关联脉络:序列到序列任务中的形式化、向计算机视觉的迁移、解决二次复杂度瓶颈的效率创新,以及可解释性进展。定义效率、表达力、可解释性三项标准,采用EEI评分框架比较21种方法,评分由单一评估者给出,假设±1分扰动范围。通过20万次确定性蒙特卡洛分析发现,平均67%-70%样本中排名变动超过一位。与秩匹配零模型结果相似,说明仅支持粗粒度层级比较,而非精细排序。综述从Bahdanau-Luong对齐机制经Transformer延伸至视觉架构,涵盖固定与学习的稀疏注意力、线性注意力、输入输出感知的精确算法(如FlashAttention)及状态空间模型(如Mamba)。还涉及归纳头、超叠加现象与注意力-状态空间模型二元性。提供结构化叙事、跨研究基准整合、五类研究空白分析及2015-2026年演进时间线。结论将注意力研究视为效率-表达力-可解释性边界的拓展,提出统一效率基准、混合架构中的学习路由、长序列泛化与可扩展机制可解释性等未来方向。
原文摘要 · Abstract (English)
Attention mechanisms have driven machine learning for a decade, from neural machine translation to language models that do general-purpose reasoning. This survey covers four connected threads: their formulation for sequence-to-sequence tasks, adaptation to computer vision, efficiency innovations that address the quadratic bottleneck, and advances in interpretability. We define three criteria: efficiency, expressiveness, and interpretability, and compare twenty-one methods using an EEI scoring framework. Scores come from a single rater with an assumed +/-1-point perturbation range. A deterministic Monte Carlo analysis with 200,000 samples shows that, under this perturbation model, rank changes of more than one position occur in 67-70% of samples on average. A rank-matched null model reproduces a similar stability profile, so the results support coarse tier-level comparisons rather than fine-grained rankings. The survey traces attention from Bahdanau-Luong alignment through the Transformer and into vision architectures. It reviews fixed and learned sparse attention, linear attention, IO-aware exact algorithms including FlashAttention, and state-space alternatives including Mamba. It also covers induction heads, superposition, and the attention-SSM duality. We further provide a structured narrative review, a benchmark synthesis with cross-study caveats, a five-problem research gap analysis, and a 2015-2026 evolution timeline. We conclude by framing attention research as an expansion of the efficiency-expressiveness-interpretability frontier and identifying future directions including unified efficiency benchmarks, learned routing for hybrid architectures, length generalization, and scalable mechanistic interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。