用舒尔分解确保状态空间网络稳定,训练更稳且高效。
A Novel Schur-Decomposition-Based Weight Projection Method for Stable State-Space Neural-Network Architectures

- 通过舒尔分解动态投影状态矩阵至最邻近稳定形式
- 在合成系统上达到与顶尖方法相当的精度和收敛速度
- 适合需严格稳定性保障的复杂动力系统建模
从数据构建动态系统的黑箱模型是机器学习中的挑战性问题,尤其当需要渐近稳定性保证时。本文提出一种基于舒尔分解的新型稳定性保障且支持反向传播的投影方案,用于线性离散时间状态空间层的状态矩阵,并提供一种替代的预分解形式。该方法将状态矩阵实舒尔分解的拟三角因子动态投影到最近的稳定形式,确保系统动态稳定且过参数化最小。在合成线性系统上的实验表明,尽管计算复杂度略有增加,但该方法在准确性和收敛速率上可媲美当前最先进的稳定系统辨识技术。此外,更低的参数量有助于在含静态非线性层的堆叠神经网络架构中实现更快收敛,同时不牺牲真实世界数据集上的准确性。结果表明,该舒尔分解投影方法为识别复杂动力系统提供了数值稳健框架,性能媲美业界前沿,且满足严格的渐近稳定性要求。
原文摘要 · Abstract (English)
Building black-box models for dynamical systems from data is a challenging problem in machine learning, especially when asymptotic stability guarantees are required. In this paper, we introduce a novel stability-ensuring and backpropagation-compatible projection scheme based on the Schur decomposition for the state matrix of linear discrete-time state-space layers, as well as an alternative pre-factorized formulation of the methodology. The proposed methods dynamically project the quasi-triangular factor of the state matrix's real Schur decomposition onto its nearest stable peer, ensuring stable dynamics with minimal overparameterization. Experiments on synthetic linear systems demonstrate that the method achieves accuracy and convergence rates comparable to those of state-of-the-art stable-system identification techniques, despite a marginal increase in computational complexity. Furthermore, the lower weight count facilitates convergence during training without sacrificing accuracy in stacked neural-network architectures with static nonlinearities targeting real-world datasets. These results suggest that the Schur-based projection provides a numerically robust framework for identifying complex dynamics on par with the State of the Art while satisfying strict asymptotic-stability requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。