arXiv:2603.16535cs.LGmath.OC2026-03被引 1

用惯性动力学加速注意力机制,收敛更快且不增加计算开销。

SympFormer: Accelerated attention blocks via Inertial Dynamics on Density Manifolds

  • 将注意力块视为密度空间上的粒子系统,引入类Nesterov惯性动力学。
  • 线性注意力近似斯坦因变分梯度流,椭圆分布保持不变。
  • 实现在相同调用次数下更快收敛,适合追求高效推理的场景。

Transformer 在自然语言处理中的成功很大程度上归功于自注意力模块。近期研究将注意力块视为相互作用的粒子系统,其均场极限对应于在配备 Wasserstein-2 型度量的概率密度空间上的交互能量泛函的梯度流。本文在此基础上,通过密度空间上的惯性 Nesterov 型动力学,提出加速注意力模块。所提架构中,标记同时携带空间(特征)与速度变量。时间离散化及加速密度动力学的近似生成了哈密顿动量注意力模块,构成新的加速注意力结构。特别地,对于线性自注意力,我们证明注意力模块近似一个使用双线性核的势能斯坦因变分梯度流。在此设定下,我们证明椭圆轮廓概率分布被加速注意力模块保持不变。本文提出可实现的基于粒子的算法,并表明所提加速注意力模块在保持调用次数不变的情况下收敛速度优于经典注意力模块。

原文摘要 · Abstract (English)

Transformers owe much of their empirical success in natural language processing to the self-attention blocks. Recent perspectives interpret attention blocks as interacting particle systems, whose mean-field limits correspond to gradient flows of interaction energy functionals on probability density spaces equipped with Wasserstein-$2$-type metrics. We extend this viewpoint by introducing accelerated attention blocks derived from inertial Nesterov-type dynamics on density spaces. In our proposed architecture, tokens carry both spatial (feature) and velocity variables. The time discretization and the approximation of accelerated density dynamics yield Hamiltonian momentum attention blocks, which constitute the proposed accelerated attention architectures. In particular, for linear self-attention, we show that the attention blocks approximate a Stein variational gradient flow, using a bilinear kernel, of a potential energy. In this setting, we prove that elliptically contoured probability distributions are preserved by the accelerated attention blocks. We present implementable particle-based algorithms and demonstrate that the proposed accelerated attention blocks converge faster than the classical attention blocks while preserving the number of oracle calls.

注意力机制加速计算扩散模型动力系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。