揭示状态空间模型的表达能力差异及其与注意力机制的关系
On the Expressiveness of State Space Models via Temporal Logics
- 用线性时序逻辑分析状态空间模型的表达能力
- 量化模型仅能表达正则语言,高精度可捕捉计数特性
- 对比了不同门控机制和与Transformer的表达力差异
我们研究了状态空间模型(SSM)的表达能力,这类模型近期被视为大语言模型中Transformer架构的潜在替代方案。基于已有工作,我们通过有限轨迹上的线性时序逻辑片段与扩展来分析SSM的表达力。结果表明,其表达能力显著依赖于底层门控机制。在固定精度算术下运行的SSM(量化模型)的表达力局限于正则语言;而具有无界精度的SSM可捕捉计数性质与非正则语言。此外,我们系统比较了不同SSM变体与已知的Transformer表达力结果,厘清了二者在表达能力上的关系。
原文摘要 · Abstract (English)
We investigate the expressive power of state space models (SSM), which have recently emerged as a potential alternative to transformer architectures in large language models. Building on recent work, we analyse SSM expressiveness through fragments and extensions of linear temporal logic over finite traces. Our results show that the expressive capabilities of SSM vary substantially depending on the underlying gating mechanism. We further distinguish between SSM operating over fixed-width arithmetic (quantised models), whose expressive power remains within regular languages, and SSM with unbounded precision, which can capture counting properties and non-regular languages. In addition, we provide a systematic comparison between these different SSM variants and known results on transformers, thereby clarifying how the two architectures relate in terms of expressive power.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。