首次分析域随机化下LQR的策略梯度收敛性,为仿真到现实迁移提供理论支持。
Policy Gradient for LQR with Domain Randomization
- 基于策略梯度优化域随机化下的线性二次调节问题
- 在系统异质性受控条件下实现全局收敛,样本复杂度可量化
- 提出退火折扣因子算法,无需初始稳定控制器,适合实际部署
域随机化(DR)通过在一系列模拟环境上训练控制器,实现从仿真到现实的迁移,目标是在真实世界中获得鲁棒性能。尽管DR在实践中广泛应用,且常采用简单的策略梯度(PG)方法求解,但其理论保证仍不清晰。本文首次对域随机化线性二次调节(LQR)问题中的策略梯度方法提供了收敛性分析。我们证明,在采样系统异质性满足特定边界条件下,PG可全局收敛至有限样本近似域随机化目标的最小值。同时,量化了实现样本均值与总体目标间小性能差距所需的样本复杂度。此外,我们提出并分析了一种折扣因子退火算法,避免了对初始联合稳定控制器的需求,该控制器可能难以找到。实验结果验证了理论发现,并指出了未来方向,包括风险敏感的域随机化形式和随机策略梯度算法。
原文摘要 · Abstract (English)
Domain randomization (DR) enables sim-to-real transfer by training controllers on a distribution of simulated environments, with the goal of achieving robust performance in the real world. Although DR is widely used in practice and is often solved using simple policy gradient (PG) methods, understanding of its theoretical guarantees remains limited. Toward addressing this gap, we provide the first convergence analysis of PG methods for domain-randomized linear quadratic regulation (LQR). We show that PG converges globally to the minimizer of a finite-sample approximation of the DR objective under suitable bounds on the heterogeneity of the sampled systems. We also quantify the sample-complexity associated with achieving a small performance gap between the sample-average and population-level objectives. Additionally, we propose and analyze a discount-factor annealing algorithm that obviates the need for an initial jointly stabilizing controller, which may be challenging to find. Empirical results support our theoretical findings and highlight promising directions for future work, including risk-sensitive DR formulations and stochastic PG algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。