arXiv:2603.04805cs.CLcs.AI2026-03被引 1

提出注意力引力场概念,揭示位置编码与模型性能的内在规律

Attention's Gravitational Field:A Power-Law Interpretation of Positional Correlation

  • 将位置编码与语义嵌入解耦,提升模型架构效率
  • 实证发现注意力分布符合幂律关系,与牛顿万有引力定律相似
  • 为注意力机制提供理论解释,适合模型优化与可解释性研究者

本文探究大语言模型中位置关系与编码的底层原理,提出注意力引力场(AGF)概念。通过将位置编码与语义嵌入解耦,优化模型架构,在准确性上优于现有编码方法。深入分析表明,AGF在内在一致性上与学习与稳定性曲线相吻合,并在经验上符合牛顿万有引力定律。本工作为理解注意力机制提供了严格的理论基础,推动了模型优化与可解释性研究的发展。

原文摘要 · Abstract (English)

This paper explores the underlying principles of positional relationships and encodings within Large Language Models (LLMs) and introduces the concept of the Attention Gravitational Field (AGF). By decoupling positional encodings from semantic embeddings, we optimize the model architecture and achieve superior accuracy compared to prevailing encoding methods. Furthermore, we provide an in-depth analysis of AGF, demonstrating its intrinsic consistency with learning and stability curves, as well as its empirical alignment with Newton's Law of Universal Gravitation. By offering a rigorous theoretical exploration of these phenomena, this work represents a significant step toward interpreting the Attention mechanism and unlocks new possibilities for future research in model optimization and interpretability.

注意力机制位置编码理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。