提出连续时间鲁棒强化学习的策略梯度方法,保障最坏情况下的性能表现。
Policy Gradient for Continuous-Time Robust Markov Decision Processes
- 基于路径和伴随法推导连续时间策略与对抗梯度
- 双循环优化器实现线性收敛与1/ε²样本复杂度
- 适用于神经微分方程动态系统,适合高安全性控制场景
鲁棒马尔可夫决策过程(RMDP)框架允许设计在最坏转移动态下仍满足性能保证的强化学习智能体。传统RMDP处理离散时间动态,近期已有高效样本的策略梯度算法。本文研究连续时间RMDP框架内的策略梯度算法,利用随机与常微分方程的路径法和伴随法推导策略梯度与对抗梯度。提出双循环优化器,在基于查询的设置中实现线性收敛,在基于样本的设置中达到$ ilde{/mathcal{O}}(rac{1}{ε^2})$样本复杂度,同时为无折扣总成本MDP框架提供了新工具。此外,提出平均场优化器作为分布式优化器,在$N$粒子近似下实现$ ilde{/mathcal{O}}(rac{1}{K})$基于查询的收敛率和$ ilde{/mathcal{O}}(rac{N^2}{ε})$样本复杂度。在具有神经常微分方程动态的连续时间RMDP上,两种优化器的有效性均得到验证。
原文摘要 · Abstract (English)
The framework of robust Markov decision processes (RMDPs) allows the design of reinforcement learning agents that satisfy performance guarantees under worst-case transition dynamics. Traditional RMDPs consider discrete-time dynamics and recently, sample-efficient policy gradient algorithms have been considered in this context. This paper investigates policy gradient algorithms within a continuous-time RMDP framework. Policy gradients and adversarial gradients are derived using pathwise and adjoint-based formulas for stochastic and ordinary differential equations. We propose double-loop optimisers to obtain linear convergence in the oracle-based setting and an $\tilde{\mathcal{O}}(\frac{1}{ε^2})$ sample complexity in the sample-based setting in an analysis which also derives novel tools for the framework of undiscounted total cost MDPs. Additionally, we propose mean-field optimisers as distributional optimisers with an $\tilde{\mathcal{O}}(\frac{1}{K})$ oracle-based convergence rate and an $\tilde{\mathcal{O}}(\frac{N^2}ε)$ sample complexity under $N$-particle approximation. The effectiveness of continuous-time policy gradient algorithms is confirmed for both optimisers on continuous-time RMDPs with neural ordinary differential equation dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。