arXiv:2506.11043stat.MLcs.LG2025-06被引 1

用现代霍普菲尔德网络统一注意力机制,实现非线性建模提升模型性能。

A Framework for Non-Linear Attention via Modern Hopfield Networks

  • 基于现代霍普菲尔德网络构建能量函数,使注意力计算成为其稳定点。
  • 能量景观中形成'上下文井',捕捉标记间的复杂关系。
  • 适用于BERT等模型的非线性注意力头设计,适合序列建模研究者。

本文提出一种沿现代霍普菲尔德网络(MNH)思路的能量函数,其稳定点对应于Vaswani等人[12]提出的注意力机制,从而统一两者框架。该能量函数的极小值形成‘上下文井’——稳定配置,封装了标记之间的上下文关系。在 $n$ 个标记嵌入上定义的能量景观中,其梯度即为注意力计算。非线性注意力机制通过增强模型对复杂关系的理解、表示学习能力以及整体效率与性能,提升了变换器模型在各类序列建模任务中的表现。粗略类比于三次样条曲线可更优地表示非线性数据,而线性模型可能不足。该方法可用于在基于Transformer的模型(如BERT[6]等)中引入非线性注意力头。

原文摘要 · Abstract (English)

In this work we propose an energy functional along the lines of Modern Hopfield Networks (MNH), the stationary points of which correspond to the attention due to Vaswani et al. [12], thus unifying both frameworks. The minima of this landscape form "context wells" - stable configurations that encapsulate the contextual relationships among tokens. A compelling picture emerges: across $n$ token embeddings an energy landscape is defined whose gradient corresponds to the attention computation. Non-linear attention mechanisms offer a means to enhance the capabilities of transformer models for various sequence modeling tasks by improving the model's understanding of complex relationships, learning of representations, and overall efficiency and performance. A rough analogy can be seen via cubic splines which offer a richer representation of non-linear data where a simpler linear model may be inadequate. This approach can be used for the introduction of non-linear heads in transformer based models such as BERT, [6], etc.

注意力机制非线性建模霍普菲尔德网络Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。