arXiv:2411.04551math.OCcs.LG2024-11被引 32

用Transformer实现任意测度间的插值,突破传统点对点映射限制。

Measure-to-measure interpolation using Transformers

  • 将Transformer视为测度到测度的映射,通过球面上粒子系统建模
  • 单个Transformer可精准匹配N组任意输入与目标测度对
  • 为多模态数据处理提供理论基础,适合研究模型表达能力者

Transformer是支撑大语言模型成功的深度神经网络架构。与传统点对点映射不同,Transformer可被看作定义在单位球面上的相互作用粒子系统的测度到测度映射:输入为提示词中词元的经验测度,其演化受连续性方程支配。事实上,Transformer不限于经验测度,原则上可处理任意输入测度。随着Transformer所处理数据类型的快速扩展,研究其作为从任意测度到另一任意测度的映射的表达能力至关重要。为此,我们在每对输入-目标测度均可通过某个传输映射匹配的最简假设下,给出了参数显式选择,使单个Transformer能精确匹配N组任意输入测度与目标测度。

原文摘要 · Abstract (English)

Transformers are deep neural network architectures that underpin the recent successes of large language models. Unlike more classical architectures that can be viewed as point-to-point maps, a Transformer acts as a measure-to-measure map implemented as specific interacting particle system on the unit sphere: the input is the empirical measure of tokens in a prompt and its evolution is governed by the continuity equation. In fact, Transformers are not limited to empirical measures and can in principle process any input measure. As the nature of data processed by Transformers is expanding rapidly, it is important to investigate their expressive power as maps from an arbitrary measure to another arbitrary measure. To that end, we provide an explicit choice of parameters that allows a single Transformer to match $N$ arbitrary input measures to $N$ arbitrary target measures, under the minimal assumption that every pair of input-target measures can be matched by some transport map.

Transformer测度映射表达能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。