提出首个可证明的线性动态系统蒸馏方法,实现高效推理且精度不随序列长度变化。
SpectraLDS: Provable Distillation for Linear Dynamical Systems
- 基于谱变换反演构建端到端凸优化框架,恢复对称LDS模型。
- 推理时每标记恒定时间与空间开销,与序列长度无关。
- 适用于语言建模等任务,提升效率同时保持预测精度。
我们提出了首个针对对称线性动态系统(LDS)的可证明识别方法,其精度保证独立于系统状态维度或有效记忆长度。该方法基于近期工作将对称LDS表示为可通过固定谱变换学习的卷积形式,展示了如何逆向这一表示,从而从谱变换中恢复LDS模型,并得到一个端到端的凸优化过程。该蒸馏方法在保持预测精度的同时,实现了每标记恒定时间与恒定空间的推理,与序列长度无关。我们在序列预测架构中评估了SpectraLDS作为组件的效果,结果表明在语言建模等任务上,精度得以保留,同时推理效率显著提升。
原文摘要 · Abstract (English)
We present the first provable method for identifying symmetric linear dynamical systems (LDS) with accuracy guarantees that are independent of the systems' state dimension or effective memory. Our approach builds upon recent work that represents symmetric LDSs as convolutions learnable via fixed spectral transformations. We show how to invert this representation, thereby recovering an LDS model from its spectral transform and yielding an end-to-end convex optimization procedure. This distillation preserves predictive accuracy while enabling constant-time and constant-space inference per token, independent of sequence length. We evaluate our method, SpectraLDS, as a component in sequence prediction architectures and demonstrate that accuracy is preserved while inference efficiency is improved on tasks such as language modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。