用连续数学解释Transformer,揭示注意力与归一化的深层机制。
A Mathematical Explanation of Transformers
- 将Transformer视为积分微分方程的离散化,从连续角度重构其结构。
- 自注意力被解释为非局部积分算子,层归一化是时变约束下的投影。
- 适合对模型原理感兴趣的理论研究者和架构设计者。
Transformer架构彻底改变了序列建模领域,并成为大语言模型突破的核心。然而,对其结构与运作的完整数学理论仍不清晰。本文提出一种新的连续框架,将Transformer严格解释为一种结构化积分微分方程的离散化。在此框架下,自注意力机制自然地表现为非局部积分算子,层归一化则被刻画为到时变约束的投影。该算子理论与变分视角为注意力、前馈层及归一化等核心组件提供了统一且可解释的基础。本方法通过在词元索引和特征维度上同时嵌入连续域,超越了以往理论分析。该框架不仅深化了理论理解,还为架构设计、分析及基于控制论的解释开辟新路径。这一新视角有助于弥合深度学习架构与连续数学建模之间的鸿沟,为可解释且理论坚实的神经网络模型发展提供基础性洞见。
原文摘要 · Abstract (English)
The Transformer architecture has revolutionized the field of sequence modeling and underpins the recent breakthroughs in large language models (LLMs). However, a comprehensive mathematical theory that explains its structure and operations remains elusive. In this work, we propose a novel continuous framework that rigorously interprets the Transformer as a discretization of a structured integro-differential equation. Within this formulation, the self-attention mechanism emerges naturally as a non-local integral operator, and layer normalization is characterized as a projection to a time-dependent constraint. This operator-theoretic and variational perspective offers a unified and interpretable foundation for understanding the architecture's core components, including attention, feedforward layers, and normalization. Our approach extends beyond previous theoretical analyses by embedding the entire Transformer operation in continuous domains for both token indices and feature dimensions. This leads to a principled and flexible framework that not only deepens on theoretical insight but also offers new directions for architecture design, analysis, and control-based interpretations. This new interpretation provides a step toward bridging the gap between deep learning architectures and continuous mathematical modeling, and contributes a foundational perspective to the ongoing development of interpretable and theoretically grounded neural network models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。