证明了选择性状态空间层比线性注意力更强大,适合长序列建模。
On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach
- 用多元多项式分析选择性状态空间层的表达能力
- 理论证明其表达力超越线性Transformer,且保持良好泛化性
- 适用于需要长序列建模的NLP与视觉任务
近期高效序列建模的发展催生了选择性状态空间层,作为Mamba架构的核心组件,在多种自然语言处理和视觉任务中表现出色。尽管Mamba在多个基准上性能达到或超过当前最优的Transformer模型,但其强大表示能力的理论基础仍不清晰。本文通过多元多项式方法研究选择性状态空间层的表达能力,证明其表达力优于线性Transformer。结果表明,Mamba在长序列建模上具备更强的表示能力,同时不牺牲泛化性能。理论结论在多个数据集上的全面实验中得到验证。
原文摘要 · Abstract (English)
Recent advances in efficient sequence modeling have introduced selective state-space layers, a key component of the Mamba architecture, which have demonstrated remarkable success in a wide range of NLP and vision tasks. While Mamba's empirical performance has matched or surpassed SoTA transformers on such diverse benchmarks, the theoretical foundations underlying its powerful representational capabilities remain less explored. In this work, we investigate the expressivity of selective state-space layers using multivariate polynomials, and prove that they surpass linear transformers in expressiveness. Consequently, our findings reveal that Mamba offers superior representational power over linear attention-based models for long sequences, while not sacrificing their generalization. Our theoretical insights are validated by a comprehensive set of empirical experiments on various datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。