通过聚焦不稳定的子空间,用更少数据快速稳定未知的线性系统。
Learning Stabilizing Policies via an Unstable Subspace Representation
- 先学习系统的不稳定子空间,再在此低维空间求解控制问题。
- 当不稳定模式远少于状态维度时,样本需求大幅降低。
- 适合对数据效率要求高的控制学习场景,如机器人、自动化系统。
我们研究了学习稳定控制策略(LTS)以稳定线性时不变(LTI)系统的问题。现有控制策略梯度方法依赖初始稳定策略,但针对未知系统设计此类策略本身即为根本难题,其难度可能与学习最优策略相当。以往工作在LTS问题上需要大量数据,且样本复杂度随环境维度呈二次增长。本文提出两阶段方法:首先学习系统的左不稳定子空间,然后在该子空间上求解一系列折扣线性二次调节器(LQR)问题,仅针对系统不稳定动态进行稳定化,从而降低控制空间的有效维度。我们为两个阶段提供了非渐近保证,并证明在不稳定子空间上操作可显著降低样本复杂度。特别地,当不稳定模式数量远小于状态维度时,该方法极大加速了稳定化过程。数值实验验证了所提方法在降低样本复杂度方面的有效性。
原文摘要 · Abstract (English)
We study the problem of learning to stabilize (LTS) a linear time-invariant (LTI) system. Policy gradient (PG) methods for control assume access to an initial stabilizing policy. However, designing such a policy for an unknown system is one of the most fundamental problems in control, and it may be as hard as learning the optimal policy itself. Existing work on the LTS problem requires large data as it scales quadratically with the ambient dimension. We propose a two-phase approach that first learns the left unstable subspace of the system and then solves a series of discounted linear quadratic regulator (LQR) problems on the learned unstable subspace, targeting to stabilize only the system's unstable dynamics and reduce the effective dimension of the control space. We provide non-asymptotic guarantees for both phases and demonstrate that operating on the unstable subspace reduces sample complexity. In particular, when the number of unstable modes is much smaller than the state dimension, our analysis reveals that LTS on the unstable subspace substantially speeds up the stabilization process. Numerical experiments are provided to support this sample complexity reduction achieved by our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。