提出新型序列模型SLiCE,兼具高效计算与最强表达力。
Structured Linear CDEs: Maximally Expressive and Parallel-in-Time Sequence Models
- 用结构化输入依赖的转移矩阵替代传统对角矩阵
- 单层即可完成复杂状态追踪,长序列泛化性能顶尖
- 适合追求高效高精度时序建模的研究者与工程师
本文提出结构化线性受控微分方程(SLiCEs),一种统一的序列建模框架。该框架采用结构化、输入相关的状态转移矩阵,在保持稠密矩阵最大表达力的同时显著降低计算成本。其涵盖现有架构如输入依赖的块对角线递归神经网络和DeltaNet的对角加低秩结构,并引入基于稀疏性和沃尔什-哈达玛变换的两种新变体。理论上证明,相较于S4D和Mamba的对角转移矩阵,使用块对角、稀疏或沃尔什-哈达玛矩阵的SLiCEs可达到稠密矩阵的最大表达力。实验表明,SLiCEs仅用单层即解决$A_5$状态追踪基准任务,在并行时间模型中取得最佳长序列泛化性能;在六个多变量时间序列分类数据集上表现媲美对数神经控制微分方程,且训练每步耗时减少20倍。
原文摘要 · Abstract (English)
This work introduces Structured Linear Controlled Differential Equations (SLiCEs), a unifying framework for sequence models with structured, input-dependent state-transition matrices that retain the maximal expressivity of dense matrices whilst being cheaper to compute. The framework encompasses existing architectures, such as input-dependent block-diagonal linear recurrent neural networks and DeltaNet's diagonal-plus-low-rank structure, as well as two novel variants based on sparsity and the Walsh-Hadamard transform. We prove that, unlike the diagonal state-transition matrices of S4D and Mamba, SLiCEs employing block-diagonal, sparse, or Walsh-Hadamard matrices match the maximal expressivity of dense matrices. Empirically, SLiCEs solve the $A_5$ state-tracking benchmark with a single layer, achieve best-in-class length generalisation on regular language tasks among parallel-in-time models, and match the performance of log neural controlled differential equations on six multivariate time-series classification datasets while cutting the average time per training step by a factor of twenty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。