arXiv:2511.08243cs.LG2025-11

为Transformer提供几何场论统一框架,揭示位置编码与注意力的深层数学本质。

A Unified Geometric Field Theory Framework for Transformers: From Manifold Embeddings to Kernel Modulation

  • 将离散位置映射到连续流形上的空间函数,实现场论视角建模。
  • 提出核积分算子与注意力机制在嵌入流形上的统一表达形式。
  • 适合对Transformer理论基础感兴趣的科研人员与深度学习架构研究者。

Transformer架构凭借自注意力机制在自然语言处理、计算机视觉和科学计算中取得巨大成功。然而,其核心组件——位置编码与注意力机制——长期缺乏统一的物理或数学解释。本文提出一个结构化理论框架,整合位置编码、核积分算子与注意力机制,进行深入理论分析。我们将离散位置(如文本词元索引、图像像素坐标)映射到连续流形上的空间函数,使Transformer层可被解释为作用于嵌入流形上的核调制算子,从而建立其场论意义上的统一描述。

原文摘要 · Abstract (English)

The Transformer architecture has achieved tremendous success in natural language processing, computer vision, and scientific computing through its self-attention mechanism. However, its core components-positional encoding and attention mechanisms-have lacked a unified physical or mathematical interpretation. This paper proposes a structural theoretical framework that integrates positional encoding, kernel integral operators, and attention mechanisms for in-depth theoretical investigation. We map discrete positions (such as text token indices and image pixel coordinates) to spatial functions on continuous manifolds, enabling a field-theoretic interpretation of Transformer layers as kernel-modulated operators acting over embedded manifolds.

Transformer几何场论注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。