揭示自注意力如何学习并泛化实体间交互关系。
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization
- 从交互视角分析自注意力,发现其能高效建模成对关系。
- 实验证明其在分布外场景下仍具良好泛化能力。
- 提出新型模块,可捕捉多实体复杂交互,适合复杂关系建模任务。
自注意力已成为现代神经网络架构的核心组件,但其理论基础仍不清晰。本文从交互实体的视角出发,涵盖多智能体强化学习中的代理或基因序列中的等位基因,证明单层线性自注意力能够高效表示、学习和泛化捕捉成对交互的函数,包括分布外情形。分析表明,在训练中观察到的交互模式多样性极低的假设下,自注意力即能作为相互作用学习器,适用于广泛的现实场景。实验验证了自注意力学习交互函数并跨群体分布及分布外场景泛化的理论洞察。基于此理论,我们提出HyperFeatureAttention,一种新神经网络模块,用于学习不同特征层级实体间的耦合关系;进一步提出HyperAttention,扩展至三元、四元乃至一般n元交互,以捕捉多实体依赖关系。
原文摘要 · Abstract (English)
Self-attention has emerged as a core component of modern neural architectures, yet its theoretical underpinnings remain elusive. In this paper, we study self-attention through the lens of interacting entities, ranging from agents in multi-agent reinforcement learning to alleles in genetic sequences, and show that a single layer linear self-attention can efficiently represent, learn, and generalize functions capturing pairwise interactions, including out-of-distribution scenarios. Our analysis reveals that self-attention acts as a mutual interaction learner under minimal assumptions on the diversity of interaction patterns observed during training, thereby encompassing a wide variety of real-world domains. In addition, we validate our theoretical insights through experiments demonstrating that self-attention learns interaction functions and generalizes across both population distributions and out-of-distribution scenarios. Building on our theories, we introduce HyperFeatureAttention, a novel neural network module designed to learn couplings of different feature-level interactions between entities. Furthermore, we propose HyperAttention, a new module that extends beyond pairwise interactions to capture multi-entity dependencies, such as three-way, four-way, or general n-way interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。