用自由概率论分析Transformer的结构动态,揭示注意力机制的本质。
A Free Probabilistic Framework for Analyzing the Transformer-based Language Models
- 将注意力建模为非交换卷积,用算子理论解析模型内部运作。
- 在自由性假设下推导出基于熵的泛化界,解释表征复杂性演化。
- 适合研究大模型理论机制的学者,提供数学视角的新洞见。
我们提出一种基于自由概率论的形式化算子理论框架,用于分析基于Transformer的语言模型。通过将词元嵌入和注意力机制建模为具有迹的W*-概率空间中的自伴算子,我们将注意力重新解释为非交换卷积,并通过自由加性卷积描述表示传播过程。这使得深层Transformer可被理解为一个谱动态系统。在自由性假设下,我们推导出基于熵的泛化边界,并对位置编码、谱演化和表征复杂性提供了新见解。本工作为大规模语言模型的结构动力学提供了严谨而理论化的视角。
原文摘要 · Abstract (English)
We present a formal operator-theoretic framework for analyzing Transformer-based language models using free probability theory. By modeling token embeddings and attention mechanisms as self-adjoint operators in a tracial \( W^* \)-probability space, we reinterpret attention as non-commutative convolution and describe representation propagation via free additive convolution. This leads to a spectral dynamic system interpretation of deep Transformers. We derive entropy-based generalization bounds under freeness assumptions and provide insight into positional encoding, spectral evolution, and representational complexity. This work offers a principled, though theoretical, perspective on structural dynamics in large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。