用自旋-浴模型解释GPT-2注意力机制,实证其预测能力。
Testing the spin-bath view of self-attention: A Hamiltonian analysis of GPT-2 Transformer
- 将注意力头建模为自旋系统,推导有效哈密顿量。
- 理论对数间隙与实际输出排名强负相关(r≈-0.70)。
- 适合作为理解大模型生成行为的物理启发工具。
Huo和Johnson提出的物理学框架将大型语言模型的注意力机制视为相互作用的两体自旋系统,为重复和偏差等现象提供了第一性原理解释。我们从一个生产级GPT-2模型中提取了完整的查询-键权重矩阵,并为每个注意力头推导出相应的有效哈密顿量。基于这些哈密顿量,我们获得了可解析的相变边界和对数间隙判据,可预测在给定上下文下哪个词应主导下一个词的概率分布。在20个事实回忆提示上对144个注意力头进行系统评估,发现理论对数间隙与模型实际的词排名间存在显著负相关(r≈-0.70,p<10⁻³)。有针对性的消融实验进一步表明,抑制与自旋-浴预测最一致的头会导致输出概率的预期偏移,证实了因果关系而非偶然关联。综上,我们的研究首次为生产级模型中的自旋-浴类比提供了强有力的经验证据。本文采用上下文场视角,提供基于物理的可解释性,并推动结合理论凝聚态物理与人工智能的新一代生成模型发展。
原文摘要 · Abstract (English)
The recently proposed physics-based framework by Huo and Johnson~\cite{huo2024capturing} models the attention mechanism of Large Language Models (LLMs) as an interacting two-body spin system, offering a first-principles explanation for phenomena like repetition and bias. Building on this hypothesis, we extract the complete Query-Key weight matrices from a production-grade GPT-2 model and derive the corresponding effective Hamiltonian for every attention head. From these Hamiltonians, we obtain analytic phase boundaries and logit gap criteria that predict which token should dominate the next-token distribution for a given context. A systematic evaluation on 144 heads across 20 factual-recall prompts reveals a strong negative correlation between the theoretical logit gaps and the model's empirical token rankings ($r\approx-0.70$, $p<10^{-3}$).Targeted ablations further show that suppressing the heads most aligned with the spin-bath predictions induces the anticipated shifts in output probabilities, confirming a causal link rather than a coincidental association. Taken together, our findings provide the first strong empirical evidence for the spin-bath analogy in a production-grade model. In this work, we utilize the context-field lens, which provides physics-grounded interpretability and motivates the development of novel generative models bridging theoretical condensed matter physics and artificial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。