arXiv:2504.00492cs.LGmath.DS2025-04被引 6

用流分解方法并行化线性注意力,提升效率与可扩展性

ParallelFlow: Parallelizing Linear Transformers via Flow Discretization

  • 将分块操作重解为系统动态的流计算,建立理论桥梁
  • 设计出复杂度更低的新算法,优于现有硬件高效算法
  • 适合关注高效序列建模与理论指导算法设计的研究者

我们提出一种基于矩阵值状态空间模型(SSMs)的理论框架,用于分析线性注意力模型。所提出的Parallel Flows方法将时间动态与实现约束解耦,可独立分析分块、并行化与信息聚合等关键组件。核心在于将分块过程重新解释为系统动态流的计算,建立起与粗糙路径理论的联系,为序列建模架构带来新洞见。作为具体应用,我们在受近期理论进展启发的广义低秩设定下分析DeltaNet。我们的方法不仅可设计出简洁高效的已有算法泛化形式,还可基于粗糙路径技术提出全新算法,其复杂度可被严格证明更低。这一双重贡献表明,严谨的理论分析既能解释现有实用方法,又能启发根本性的新计算范式。

原文摘要 · Abstract (English)

We present a theoretical framework for analyzing linear attention models through matrix-valued state space models (SSMs). Our approach, Parallel Flows, provides a perspective that systematically decouples temporal dynamics from implementation constraints, enabling independent analysis of critical algorithmic components: chunking, parallelization, and information aggregation. Central to this framework is the reinterpretation of chunking procedures as computations of the flows governing system dynamics. This connection establishes a bridge to mathematical tools from rough path theory, opening the door to new insights into sequence modeling architectures. As a concrete application, we analyze DeltaNet in a generalized low-rank setting motivated by recent theoretical advances. Our methods allow us to design simple, streamlined generalizations of hardware-efficient algorithms present in the literature, and to provide completely different ones, inspired by rough paths techniques, with provably lower complexity. This dual contribution demonstrates how principled theoretical analysis can both explain existing practical methods and inspire fundamentally new computational approaches.

线性注意力并行计算状态空间模型算法优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。