用凸学习+蒸馏,从非线性系统中提取高效线性模型。
Spectral Distillation: From Nonlinear Dynamics to Linear State-Space Models

- 先用凸方法学隐式谱预测器,再蒸馏成显式线性系统。
- 平均误差由可忽略的蒸馏项和最优观测器复杂度决定。
- 无需解非凸问题,适合建模复杂动态系统。
能否通过紧凑的线性状态空间表示来学习未知的非线性动力系统,而无需直接求解非凸系统辨识问题?我们提出一个可证明的完整流程。从对未知非线性系统的观测出发,首先使用观察谱滤波(OSF)这一凸方法学习一个隐式谱预测器,其性能可媲美最优线性观测器。随后通过谱到线性状态空间模型(LDS)的蒸馏,将该预测器转化为显式递归线性系统。主要定理表明,蒸馏后LDS的平均预测误差分解为指数级小的蒸馏项和受最优观测器吕恩伯格复杂度控制的学习项。该保证为维度无关:依赖于观测器复杂度而非表示非线性系统所需的潜在维度。据我们所知,这是首个端到端可证明的方法,通过凸学习结合可证明蒸馏,提取非线性动力学的最优线性状态空间表示。在线性LDS基准和MuJoCo行为克隆任务上的实验表明,训练-蒸馏流程生成的紧凑LDS预测器,在性能上达到或超越直接训练的基线。
原文摘要 · Abstract (English)
Can nonlinear dynamical systems be learned through a compact linear state-space representation, without directly solving a non-convex system-identification problem? We give a provable pipeline for doing so. Starting from observations of an unknown nonlinear dynamical system, we first learn an implicit spectral predictor using Observation Spectral Filtering (OSF), a convex method that competes with the best linear observer for the system. We then apply spectral-to-LDS distillation to convert this predictor into an explicit recurrent linear dynamical system. Our main theorem shows that the average prediction error of the distilled LDS decomposes into an exponentially-small distillation term and the OSF learning term governed by the Luenberger complexity of the best observer. The guarantee is dimension-free: it depends on observer complexity rather than on the latent dimension needed to represent the nonlinear system. To our knowledge, this yields the first end-to-end provable method for extracting a best-in-hindsight LDS representation of nonlinear dynamics through convex learning followed by provable distillation. Experiments on linear LDS benchmarks and MuJoCo behavior cloning show that the train-then-distill pipeline produces compact LDS predictors that match or outperform directly trained baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。