arXiv:2410.06833cs.LGmath.AP2024-10被引 53

揭示Transformer注意力机制中粒子集群的动态亚稳态现象

Dynamic metastability in the self-attention model

  • 将自注意力模型视为球面上的粒子系统,分析其演化机制
  • 证明粒子虽终会聚为单簇,但会在多簇状态停留指数级长时间
  • 发现能量在缩放时间下呈阶梯状上升,类比神经网络训练轨迹

我们研究自注意力模型——一个定义在单位球面上的相互作用粒子系统,作为Transformer架构的简化模型。证明了[GLPR23]中提出的动态亚稳态现象:尽管粒子在无限时间内最终坍缩为单一簇,但在指数级长的时间内会持续停留在多个簇的配置附近。通过将系统解释为梯度流,我们将其与Otto和Reznikoff [OR07]提出的关于粗化和Allen-Cahn方程的梯度流慢运动框架相联系。最后,我们在亚稳态期之后探查系统动力学,发现经适当时间重标度后,能量在有限时间内达到全局最大值,呈现阶梯状轮廓,轨迹表现出类似鞍点到鞍点的行为,这与近期关于两层神经网络训练动力学的研究结果相似。

原文摘要 · Abstract (English)

We consider the self-attention model - an interacting particle system on the unit sphere, which serves as a toy model for Transformers, the deep neural network architecture behind the recent successes of large language models. We prove the appearance of dynamic metastability conjectured in [GLPR23] - although particles collapse to a single cluster in infinite time, they remain trapped near a configuration of several clusters for an exponentially long period of time. By leveraging a gradient flow interpretation of the system, we also connect our result to an overarching framework of slow motion of gradient flows proposed by Otto and Reznikoff [OR07] in the context of coarsening and the Allen-Cahn equation. We finally probe the dynamics beyond the exponentially long period of metastability, and illustrate that, under an appropriate time-rescaling, the energy reaches its global maximum in finite time and has a staircase profile, with trajectories manifesting saddle-to-saddle-like behavior, reminiscent of recent works in the analysis of training dynamics via gradient descent for two-layer neural networks.

Transformer注意力机制动力系统亚稳态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。