arXiv:2509.04154cs.LGcs.AI2025-09中稿 · ICML被引 2

将自注意力视为鲁棒状态估计,提升长序列泛化能力

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation

  • 用随机微分方程建模令牌为噪声观测,按一致性动态分配注意力
  • 在语言建模上比RoPE更低困惑度,且零样本外推到更长上下文仍稳定
  • 揭示旋转位置编码与记忆衰减的动态机制,适合关注模型可解释性者

我们提出鲁棒滤波注意力(RFA),将自注意力建模为鲁棒状态估计。每个令牌被视为由线性随机微分方程(SDE)驱动的潜在轨迹的噪声观测,注意力权重基于该模型下的一致性确定,而非静态特征相似性。在各向同性噪声和衰减假设下,RFA 的计算复杂度与标准注意力相当。在语言建模基准测试中,RFA 在训练窗口内实现比 RoPE 更低的困惑度,并在零样本外推至更长上下文时保持稳定。该框架还为标准位置编码提供了动态解释,将旋转嵌入与近期偏差关联到随机动力学引起的传输和不确定性传播。

原文摘要 · Abstract (English)

We introduce Robust Filter Attention (RFA), a formulation of self-attention as a robust state estimator. Each token is treated as a noisy observation of a latent trajectory governed by a linear stochastic differential equation (SDE), and attention weights are determined by consistency under this model rather than static feature similarity. Under isotropic noise and decay assumptions, RFA matches the computational complexity of standard attention. On language modeling benchmarks, RFA achieves lower perplexity than RoPE within the training window while remaining stable under zero-shot extrapolation to longer contexts. The framework also provides a dynamical interpretation of standard positional mechanisms, connecting rotational embeddings and recency biases to transport and uncertainty propagation induced by stochastic dynamics.

自注意力状态估计位置编码长序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。