用数学方程解析Transformer深层动态,揭示注意力机制如何改变数据分布。
A Unified Perspective on the Dynamics of Deep Transformers
- 将Transformer建模为概率测度的演化方程,统一分析多种注意力机制。
- 证明在高斯初始条件下,数据分布保持高斯性,可解析追踪各层变化。
- 发现深层中数据出现聚类现象,为理解模型行为提供理论依据。
Transformer在多数机器学习任务中表现卓越,其核心是将数据表示为向量序列(标记),并通过注意力函数学习标记间的依赖关系。然而,注意力在多层中的迭代应用引发复杂动态,尚不完全清楚。本文将输入序列视为概率测度,建立基于非线性概率测度的Vlasov型方程(Transformer PDE),用于刻画其演化。针对紧支集初始数据,证明该方程适定,并作为交互粒子系统的均场极限,推广至多头注意力、L2注意力、Sinkhorn注意力、Sigmoid注意力及掩码注意力等变体,利用条件Wasserstein框架实现统一分析。进一步首次研究非紧支集初始条件,聚焦高斯初始数据,证明不同注意力下变压器PDE保持高斯测度空间不变,从而可理论与数值分析典型行为。该分析揭示了深层Transformer中数据各向异性的演化过程,尤其发现一种与非归一化离散情形相似的聚类现象。
原文摘要 · Abstract (English)
Transformers, which are state-of-the-art in most machine learning tasks, represent the data as sequences of vectors called tokens. This representation is then exploited by the attention function, which learns dependencies between tokens and is key to the success of Transformers. However, the iterative application of attention across layers induces complex dynamics that remain to be fully understood. To analyze these dynamics, we identify each input sequence with a probability measure and model its evolution as a Vlasov equation called Transformer PDE, whose velocity field is non-linear in the probability measure. Our first set of contributions focuses on compactly supported initial data. We show the Transformer PDE is well-posed and is the mean-field limit of an interacting particle system, thus generalizing and extending previous analysis to several variants of self-attention: multi-head attention, L2 attention, Sinkhorn attention, Sigmoid attention, and masked attention--leveraging a conditional Wasserstein framework. In a second set of contributions, we are the first to study non-compactly supported initial conditions, by focusing on Gaussian initial data. Again for different types of attention, we show that the Transformer PDE preserves the space of Gaussian measures, which allows us to analyze the Gaussian case theoretically and numerically to identify typical behaviors. This Gaussian analysis captures the evolution of data anisotropy through a deep Transformer. In particular, we highlight a clustering phenomenon that parallels previous results in the non-normalized discrete case.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。