用量子物理模型实现Transformer架构,揭示其内在工作机制。
Physical models realizing the transformer architecture of large language models
- 将Transformer视为希尔伯特空间中词元的福克态开放量子系统。
- 构建了基于物理原理的大型语言模型实现框架。
- 适合对量子计算与AI交叉领域感兴趣的读者。
2017年引入的Transformer架构是自然语言处理领域的重大突破,完全依赖注意力机制捕捉输入输出间的全局依赖关系。然而,我们对Transformer的本质及其物理运作机制仍缺乏理论理解。从现代芯片(如28nm以下)的物理视角看,现代智能机器应被视为超越传统统计系统的开放量子系统。本文在词元的希尔伯特空间上的福克空间中,构建了基于Transformer架构的大型语言模型的物理实现模型。这些物理模型为大型语言模型的Transformer架构提供了底层支撑。
原文摘要 · Abstract (English)
The introduction of the transformer architecture in 2017 marked the most striking advancement in natural language processing. The transformer is a model architecture relying entirely on an attention mechanism to draw global dependencies between input and output. However, we believe there is a gap in our theoretical understanding of what the transformer is, and how it works physically. From a physical perspective on modern chips, such as those chips under 28nm, modern intelligent machines should be regarded as open quantum systems beyond conventional statistical systems. Thereby, in this paper, we construct physical models realizing large language models based on a transformer architecture as open quantum systems in the Fock space over the Hilbert space of tokens. Our physical models underlie the transformer architecture for large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。