arXiv:2502.14210math.OCcs.LG2025-02被引 1

新算法无需初始稳定策略,仍保持高效样本复杂度。

Sample Complexity of Linear Quadratic Regulator Without Initial Stability

  • 基于滚动时域设计,避免依赖双点梯度估计
  • 理论证明样本复杂度与最优阶数一致
  • 适合无初始稳定性保障的强化学习控制场景

受REINFORCE启发,我们提出一种新型滚动时域算法解决未知动态的线性二次调节器(LQR)问题。与以往方法不同,该算法不依赖双点梯度估计,同时保持相同阶数的样本复杂度。此外,它消除了对初始策略必须稳定的限制,显著拓展了适用范围。通过在黎曼距离下对Riccati算子收缩性的精细分析,我们实现了更优的误差传播控制,进一步降低样本复杂度并强化收敛性保证。

原文摘要 · Abstract (English)

Inspired by REINFORCE, we introduce a novel receding-horizon algorithm for the Linear Quadratic Regulator (LQR) problem with unknown dynamics. Unlike prior methods, our algorithm avoids reliance on two-point gradient estimates while maintaining the same order of sample complexity. Furthermore, it eliminates the restrictive requirement of starting with a stable initial policy, broadening its applicability. Beyond these improvements, we introduce a refined analysis of error propagation through the contraction of the Riccati operator under the Riemannian distance. This refinement leads to a better sample complexity and ensures improved convergence guarantees.

强化学习LQR样本效率控制理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。