arXiv:2602.01468cs.LGstat.ML2026-02

揭示门控注意力的统计优势:比多头注意力更省数据

A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts

  • 将门控注意力重构为分层专家混合模型
  • 门控注意力仅需多项式样本量即可准确估计专家
  • 适合关注模型效率与理论分析的研究者

自注意力机制通过捕捉长程依赖关系,极大推动了Transformer架构的成功。为提升性能,近期提出了一种在多头自注意力中引入门控机制的门控注意力模型,实证表明其能增强低秩映射的表达能力,并消除注意力塌陷现象。然而,该方法的理论机制尚不清晰。本文严谨证明:门控注意力矩阵中的每个元素均可表示为分层专家混合形式。通过将学习转化为专家估计问题,我们发现门控注意力比多头自注意力更具样本效率——前者仅需多项式数量的数据点即可实现相同估计误差,而后者需指数级数据量。此外,我们的分析还从理论上解释了为何在缩放点积注意力或值映射输出处放置门控器效果更优。

原文摘要 · Abstract (English)

Self-attention has greatly contributed to the success of the widely used Transformer architecture by enabling learning from data with long-range dependencies. In an effort to improve performance, a gated attention model that leverages a gating mechanism within the multi-head self-attention has recently been proposed as a promising alternative. Gated attention has been empirically demonstrated to increase the expressiveness of low-rank mapping in standard attention and even to eliminate the attention sink phenomenon. Despite its efficacy, a clear theoretical understanding of gated attention's benefits remains lacking in the literature. To close this gap, we rigorously show that each entry in a gated attention matrix or a multi-head self-attention matrix can be written as a hierarchical mixture of experts. By recasting learning as an expert estimation problem, we demonstrate that gated attention is more sample-efficient than multi-head self-attention. In particular, while the former needs only a polynomial number of data points to estimate an expert, the latter requires exponentially many data points to achieve the same estimation error. Furthermore, our analysis also provides a theoretical justification for why gated attention yields higher performance when a gate is placed at the output of the scaled dot product attention or the value map rather than at other positions in the multi-head self-attention architecture.

注意力机制理论分析专家混合样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。