复杂参数化让状态空间模型更高效,理论证明其优势显著。
Provable Benefits of Complex Parameterizations for Structured State Space Models
- 用复数参数可更高效表达映射关系
- 实数模型需指数级大参数才能实现同等效果
- 适合研究高效神经网络结构的学者
结构化状态空间模型(SSMs)是S4、Mamba等先进神经网络的核心,通常采用对角结构。与常规神经网络不同,SSMs常使用复数参数。本文首次从理论上揭示复数参数的优势:第一,复数对角SSM仅需较低维度即可表达实数SSM的所有映射,而反向则需高得多的维度;第二,即使实数模型维度足够,也需指数级大的参数值才能实现,难以学习;而复数模型可用中等参数值完成相同任务。实验验证了理论结论,并提示可扩展理论以包含选择性机制——一种带来当前最优性能的新架构特征。
原文摘要 · Abstract (English)
Structured state space models (SSMs), the core engine behind prominent neural networks such as S4 and Mamba, are linear dynamical systems adhering to a specified structure, most notably diagonal. In contrast to typical neural network modules, whose parameterizations are real, SSMs often use complex parameterizations. Theoretically explaining the benefits of complex parameterizations for SSMs is an open problem. The current paper takes a step towards its resolution, by establishing formal gaps between real and complex diagonal SSMs. Firstly, we prove that while a moderate dimension suffices in order for a complex SSM to express all mappings of a real SSM, a much higher dimension is needed for a real SSM to express mappings of a complex SSM. Secondly, we prove that even if the dimension of a real SSM is high enough to express a given mapping, typically, doing so requires the parameters of the real SSM to hold exponentially large values, which cannot be learned in practice. In contrast, a complex SSM can express any given mapping with moderate parameter values. Experiments corroborate our theory, and suggest a potential extension of the theory that accounts for selectivity, a new architectural feature yielding state of the art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。