新算法无需初始稳定策略,仍保持高效样本复杂度。
Sample Complexity of Linear Quadratic Regulator Without Initial Stability
- 基于滚动时域设计,避免依赖双点梯度估计
- 理论证明样本复杂度与最优阶数一致
- 适合无初始稳定性保障的强化学习控制场景
受REINFORCE启发,我们提出一种新型滚动时域算法解决未知动态的线性二次调节器(LQR)问题。与以往方法不同,该算法不依赖双点梯度估计,同时保持相同阶数的样本复杂度。此外,它消除了对初始策略必须稳定的限制,显著拓展了适用范围。通过在黎曼距离下对Riccati算子收缩性的精细分析,我们实现了更优的误差传播控制,进一步降低样本复杂度并强化收敛性保证。
原文摘要 · Abstract (English)
Inspired by REINFORCE, we introduce a novel receding-horizon algorithm for the Linear Quadratic Regulator (LQR) problem with unknown dynamics. Unlike prior methods, our algorithm avoids reliance on two-point gradient estimates while maintaining the same order of sample complexity. Furthermore, it eliminates the restrictive requirement of starting with a stable initial policy, broadening its applicability. Beyond these improvements, we introduce a refined analysis of error propagation through the contraction of the Riccati operator under the Riemannian distance. This refinement leads to a better sample complexity and ensures improved convergence guarantees.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。