用控制理论分析Mamba类模型的稳定性,提出可训练的正则化方法。
Regularity and Stability Properties of Selective SSMs with Discontinuous Gating
- 通过耗散性与输入-状态稳定理论,分离选择信号与输入建模
- 在7个时序数据集上使核心LMI违规减少92%,误差增加不足0.018%
- 适用于需要高稳定性的长序列建模任务,如金融、医疗时间序列
选择性状态空间模型(SSM)如Mamba已成为长序列建模的核心。然而其稳定性尚不明确:状态空间系数由依赖标记的门控信号在线调制,导致递推既非线性时不变也非经典非线性。本文通过无源性、耗散性及输入-状态稳定性(ISS)研究连续时间选择性SSM,显式分离选择信号 $x(ullet)$ 与驱动输入 $u(ullet)$。获得四项结果:严格耗散性下的指数遗忘;冻结选择子系统对应的局部AUC_{ ext{loc}}二次储能;参数化LMI结合通用核约束与“不可逆遗忘”;在可接受选择调度下全局一致的ISS充分条件。进一步通过推导采样块LMI桥接实际应用,作为Mamba选择扫描核心的可微训练正则化器。在七个标准时间序列数据集和四个预测跨度上,该正则化器在28/28组对比中将采样核心LMI违反降低约92%,清洁均方误差代价低于0.018%。同时改善注入扰动下的内部被动性和状态范数诊断。研究将经典控制工具转化为选择性SSM的可验证结构与训练准则,同时如实界定其向深层选择扫描架构的保证迁移范围。
原文摘要 · Abstract (English)
Selective State-Space Models (SSMs) such as Mamba have become central to long-sequence modeling. Still, their stability is poorly understood: their state-space coefficients are modulated online by a token-dependent gating signal, making the recurrence neither linear time-invariant nor classically nonlinear. We study continuous-time selective SSMs through passivity, dissipativity, and Input-to-State Stability (ISS), explicitly separating the selection signal $x(\cdot)$ from the driving input $u(\cdot)$. We obtain four results: exponential forgetting under strict dissipativity; a canonical $\mathrm{AUC}_{\mathrm{loc}}$ quadratic storage for the frozen-selection subsystem that accommodates discontinuous gating; a parametric LMI together with universal kernel constraints and "irreversible forgetting" under universal quadratic storage; and sufficient conditions for global ISS uniformly over admissible selection schedules. We then bridge to practice by deriving a sampled block LMI for the Mamba selective-scan core, which is used as a differentiable training-time regularizer. Across seven standard time-series datasets and four prediction horizons, the regularizer reduces sampled Mamba-core LMI violations by roughly $92\%$ in $28/28$ pairs at a clean-MSE cost of less than $0.018\%$. It improves internal Mamba passivity and state-norm diagnostics under injected perturbations. Our results turn classical control-theoretic tools into verifiable structural and training criteria for selective SSMs, while honestly scoping which guarantees transfer to a deep selective-scan architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。