揭示Transformer在测度映射下的数学本质,证明其可逼近非局部传输系统演化过程。
Transformers through the lens of support-preserving maps between measures
- 从测度映射角度刻画Transformer的数学结构
- 证明无限深度测度Transformer可逼近维拉索夫方程解流
- 适用于理解大模型泛化与连续动态系统的建模
Transformer是定义‘上下文映射’的深层架构,能够基于给定的上下文令牌集(如NLP中的提示或视觉Transformer中的图像块)预测新令牌。此前工作研究了这些架构处理任意数量上下文令牌的能力。为统一分析其表达能力,我们考虑映射以概率分布形式表示上下文,有限令牌数下该分布为离散型。将神经网络建模为概率测度上的映射具有多重应用,如研究Wasserstein正则性、证明泛化界以及对粒子相互作用动力学进行均场极限分析。本文研究何种测度间映射可由Transformer实现:通过前推(push forward)表示的上下文映射,其特征是保持支撑集基数且弗雷歇导数的规则部分一致连续。这些性质既涵盖Transformer,又使其能统一逼近任意连续的上下文映射。此外,我们证明均场极限下交互粒子系统的柯西问题对应的维拉索夫方程解流满足上述条件,因而可被Transformer逼近;反之,测度自注意力具备相应性质,使无穷深度均场测度变压器可等同于维拉索夫流。
原文摘要 · Abstract (English)
Transformers are deep architectures that define ``in-context maps'' which enable predicting new tokens based on a given set of tokens (such as a prompt in NLP applications or a set of patches for a vision transformer). In previous work, we studied the ability of these architectures to handle an arbitrarily large number of context tokens. To mathematically, uniformly analyze their expressivity, we considered the case that the mappings are conditioned on a context represented by a probability distribution which becomes discrete for a finite number of tokens. Modeling neural networks as maps on probability measures has multiple applications, such as studying Wasserstein regularity, proving generalization bounds and doing a mean-field limit analysis of the dynamics of interacting particles as they go through the network. In this work, we study the question what kind of maps between measures are transformers. We fully characterize the properties of maps between measures that enable these to be represented in terms of in-context maps via a push forward. On the one hand, these include transformers; on the other hand, transformers universally approximate representations with any continuous in-context map. These properties are preserving the cardinality of support and that the regular part of their Fréchet derivative is uniformly continuous. Moreover, we show that the solution map of the Vlasov equation, which is of nonlocal transport type, for interacting particle systems in the mean-field regime for the Cauchy problem satisfies the conditions on the one hand and, hence, can be approximated by a transformer; on the other hand, we prove that the measure-theoretic self-attention has the properties that ensure that the infinite depth, mean-field measure-theoretic transformer can be identified with a Vlasov flow.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。