arXiv:2603.07437cs.LGcs.SY2026-03

通过预测累计成本学习状态表示,实现无限时域LQG控制的近优解。

Cost-Driven Representation Learning for Linear Quadratic Gaussian Control: Part II

  • 基于累计成本预测构建隐空间动力学模型,实现成本驱动的状态表征学习。
  • 在无限时域时不变LQG控制中,证明了有限样本下可找到近优表示函数与控制器。
  • 方法借鉴MuZero思想,且首次证明相关随机过程的持续激励性,具理论价值。

我们研究从部分且可能高维观测中学习控制的状态表示问题。通过成本驱动的状态表示学习,我们在隐状态空间中学习动态模型,以预测累计成本。特别地,本文建立了在无限时域时不变线性二次高斯(LQG)控制中,使用所学隐模型寻找近优表示函数和近优控制器的有限样本保证。研究了两种成本驱动表示学习方法:一种显式学习隐状态转移函数,另一种隐式学习动力学,通过预测累计成本实现,后者与近期强化学习突破性方法MuZero高度相似。本文的关键技术贡献在于证明了新出现的随机过程在二次回归分析中的持续激励性,该结果可能具有独立研究意义。

原文摘要 · Abstract (English)

We study the problem of state representation learning for control from partial and potentially high-dimensional observations. We approach this problem via cost-driven state representation learning, in which we learn a dynamical model in a latent state space by predicting cumulative costs. In particular, we establish finite-sample guarantees on finding a near-optimal representation function and a near-optimal controller using the learned latent model for infinite-horizon time-invariant Linear Quadratic Gaussian (LQG) control. We study two approaches to cost-driven representation learning, which differ in whether the transition function of the latent state is learned explicitly or implicitly. The first approach has also been investigated in Part I of this work, for finite-horizon time-varying LQG control. The second approach closely resembles MuZero, a recent breakthrough in empirical reinforcement learning, in that it learns latent dynamics implicitly by predicting cumulative costs. A key technical contribution of this Part II is to prove persistency of excitation for a new stochastic process that arises from the analysis of quadratic regression in our approach, and may be of independent interest.

控制表示学习强化学习理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。