提出精确流线性注意力,消除离散化误差,提升模型稳定性和性能。
Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
- 基于连续时间系统重构更新机制,用闭式解替代欧拉离散化。
- 在语言建模和合成基准上,困惑度降低,下游性能超越现有基线。
- 无额外参数,保持线性复杂度与并行计算优势,适合长序列任务。
本文提出精确流线性注意力(EFLA),一种对增量规则线性注意力的精确流公式。我们证明增量规则更新可视为底层连续时间系统的显式欧拉离散化。EFLA以精确闭式流替换一阶更新。通过利用动态矩阵的秩-1结构,矩阵指数与输入积分均简化为一个更新步骤,同时保留增量规则线性注意力的代数结构、参数量、线性时间复杂度及分块并行性。该注意力机制在不引入额外参数的前提下,消除了增量规则动态中的欧拉离散化误差。在鲁棒性测试、语言建模范式及MAD合成基准上的实验表明,EFLA在噪声与高能输入下更具稳定性,降低困惑度,并优于SSM与欧拉风格基线的下游表现。结果确立了精确流积分作为增量规则线性注意力的一种原则性且可扩展的更新机制。
原文摘要 · Abstract (English)
In this paper, we introduce Exact Flow Linear Attention~(EFLA), an exact-flow formulation of delta-rule linear attention. We show that the delta-rule update can be interpreted as an explicit Euler discretization of an underlying continuous-time system. EFLA replaces this first-order update with the exact closed-form flow. By exploiting the rank-1 structure of the dynamics matrix, both the matrix exponential and the input integral collapse to a simple update that preserves delta-rule linear attention's algebraic structure, parameter count, linear-time complexity, and chunkwise parallelism. This attention mechanism removes the Euler discretization error of the delta-rule dynamics without introducing additional parameters. Experiments on robustness tests, language modeling benchmarks, and the MAD synthetic benchmark show that EFLA improves stability under corrupted and high-energy inputs, reduces perplexity, and achieves stronger downstream performance compared to SSM and Euler-style baselines. These results establish exact-flow integration as a principled and scalable update mechanism for delta-rule linear attention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。