arXiv:2603.13587math.OCcs.LG2026-03

用控制理论解析状态空间模型训练机制

State-space models through the lens of ensemble control

  • 将SSM训练建模为群体最优控制问题,统一处理输入依赖系统
  • 推导出PMP必要条件,证明算法子序列收敛与全局最优性
  • 为SSM训练提供可解释的控制论视角,适合研究者参考

状态空间模型(SSMs)在序列建模中表现优异,但其训练动态缺乏严谨的理论理解。本文将SSM训练形式化为一个群体最优控制问题,其中共享控制律协调一组输入依赖的动力系统。我们推导了该群体控制公式的庞特里亚金最大值原理(PMP),给出了最优性的必要条件。基于这些条件,提出一种基于逐次逼近方法的算法,并证明该迭代方案沿子序列收敛,同时建立了全局最优的充分条件。所提出的框架为SSM训练提供了控制论视角。

原文摘要 · Abstract (English)

State-space models (SSMs) are effective architectures for sequential modeling, but a rigorous theoretical understanding of their training dynamics is still lacking. In this work, we formulate the training of SSMs as an ensemble optimal control problem, where a shared control law governs a population of input-dependent dynamical systems. We derive Pontryagin's maximum principle (PMP) for this ensemble control formulation, providing necessary conditions for optimality. Motivated by these conditions, we introduce an algorithm based on the method of successive approximations. We prove convergence of this iterative scheme along a subsequence and establish sufficient conditions for global optimality. The resulting framework provides a control-theoretic perspective on SSM training.

状态空间模型控制理论优化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。