gating让注意力模型具备曲率表达能力,突破传统线性限制
Gating Enables Curvature: A Geometric Expressivity Gap in Attention

- 用统计流形几何分析注意力机制,揭示门控引入非平坦曲率
- 门控模型在非线性任务上表现更好,曲率随深度累积增强
- 适合研究模型几何表达能力或改进复杂推理任务的读者
乘法门控广泛应用于神经网络架构,并被引入注意力层以提升大语言模型的性能与训练稳定性。尽管门控注意力表现优异,其数学机理仍不清晰。本文通过将注意力输出建模为高斯分布的均值参数,分析其诱导的Fisher-Rao几何结构。结果表明,无门控注意力因仿射结构限制于内在平坦的统计流形,而乘法门控可实现非平坦几何,包括无门控情况下无法达到的正曲率流形。该发现确立了无门控与门控注意力之间的几何表达能力差距。实验显示,门控模型在需非线性决策边界的任务中具有更高表示曲率和更优性能,而在线性任务中无显著优势。此外,我们识别出一种结构化状态,在该状态下曲率在堆叠中累积,产生系统性的深度放大效应。
原文摘要 · Abstract (English)
Multiplicative gating is widely used in neural architectures and has recently been applied to attention layers to improve performance and training stability in large language models. Despite the success of gated attention, the mathematical implications of gated attention mechanisms remain poorly understood. We study attention through the geometry of its representations by modeling outputs as mean parameters of Gaussian distributions and analyzing the induced Fisher--Rao geometry. We show that ungated attention operator is restricted to intrinsically flat statistical manifolds due to its affine structure, while multiplicative gating enables non-flat geometries, including positively curved manifolds that are unattainable in the ungated setting. These results establish a geometric expressivity gap between ungated and gated attention. Empirically, we show that gated models exhibit higher representation curvature and improved performance on tasks requiring nonlinear decision boundaries whereas they provide no consistent advantage on tasks with linear decision boundaries. Furthermore, we identify a structured regime in which curvature accumulates under composition, yielding a systematic depth amplification effect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。