用Transformer注意力头构建量子场论,实现非高斯场统计与可调控关联。
Neural Network Quantum Field Theory from Transformer Architectures
- 通过注意力权重定义场论的多点关联函数,引入随机参数生成非高斯统计。
- 四点关联函数中出现有限的“独立性破坏”贡献,源于查询-键权重协方差。
- 多头求平均后非高斯关联按1/N_h衰减,大头数下趋于高斯理论。
我们提出一种基于Transformer注意力头的神经网络构造方法,用于欧几里得标量量子场论,通过在神经网络量子场论(NN-QFT)框架下对随机网络参数取平均来定义n点关联函数。单个注意力头中,共享的随机Softmax权重耦合不同宽度坐标,导致非高斯场统计,并在无限宽极限d_k→∞下依然存在。我们以注意力权重形式计算两点函数,表明可通过随机特征标记嵌入设计满足欧几里得不变性的核函数。随后分析连通四点函数,识别出一个“独立性破坏”项,其可表示为查询-键权重的协方差,在无限宽度下仍保持有限。最后证明,将多个独立注意力头以标准1/N_h归一化求和,可使连通非高斯关联函数按1/N_h衰减,从而在大量头的极限下得到高斯型NN-QFT。
原文摘要 · Abstract (English)
We propose a neural-network construction of Euclidean scalar quantum field theories from transformer attention heads, defining $n$-point correlators by averaging over random network parameters in the NN-QFT framework. For a single attention head, shared random softmax weights couple different width coordinates and induce non-Gaussian field statistics that persist in the infinite-width limit $d_k\to\infty$. We compute the two-point function in an attention-weight representation and show how Euclidean-invariant kernels can be engineered via random-feature token embeddings. We then analyze the connected four-point function and identify an "independence-breaking" contribution, expressible as a covariance over query-key weights, which remains finite at infinite width. Finally, we show that summing many independent heads with standard $1/N_h$ normalization suppresses connected non-Gaussian correlators as $1/N_h$, yielding a Gaussian NN-QFT in the large-head limit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。