arXiv:2409.10489cs.LGcs.AI2024-09被引 6

提出一种新型混合架构,实现高效长序列建模。

Flash STU: Fast Spectral Transform Units

  • 融合频谱状态空间与滑动窗口注意力,提升模型表达能力。
  • 在固定参数下,优于Transformer、S4和Mamba-2等主流模型。
  • 支持百亿参数规模,保持近线性计算复杂度,适合大规模语言建模。

近期的状态空间模型架构在高效序列建模方面展现出巨大潜力,但如何平衡计算效率与模型表达力仍是挑战。我们提出Flash STU架构,一种将频谱状态空间层与滑动窗口注意力交替结合的混合模型,可在保持近线性时间复杂度的同时,支持语言建模中数十亿参数的扩展。我们在多种序列预测任务上评估了Flash STU及其变体,包括线性动态系统、机器人控制和语言建模。结果显示,在固定参数预算下,Flash STU始终优于Transformer及其他先进状态空间模型(如S4和Mamba-2)。

原文摘要 · Abstract (English)

Recent advances in state-space model architectures have shown great promise for efficient sequence modeling, but challenges remain in balancing computational efficiency with model expressiveness. We propose the Flash STU architecture, a hybrid model that interleaves spectral state space model layers with sliding window attention, enabling scalability to billions of parameters for language modeling while maintaining a near-linear time complexity. We evaluate the Flash STU and its variants on diverse sequence prediction tasks, including linear dynamical systems, robotics control, and language modeling. We find that, given a fixed parameter budget, the Flash STU architecture consistently outperforms the Transformer and other leading state-space models such as S4 and Mamba-2.

状态空间模型序列建模高效架构语言建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。