提出离散采样随机策略的理论框架,解决连续时间强化学习执行难题。
Accuracy of Discretely Sampled Stochastic Policies in Continuous-time Reinforcement Learning
- 在离散时间点采样随机策略并作为分段常数控制执行
- 证明采样率趋近零时状态过程弱收敛,收敛速度达一阶最优
- 为策略评估与梯度估计提供偏差方差分析,适合研究连续控制理论者
随机策略(又称松弛控制)广泛应用于连续时间强化学习算法中。然而,在连续时间环境中执行随机策略并评估其性能仍是开放挑战。本文提出并严格分析了一种策略执行框架:在离散时间点从随机策略中采样动作,并将其作为分段常数控制实施。我们证明当采样网格大小趋于零时,受控状态过程弱收敛到由随机策略系数平均后的动力学系统。基于系数的正则性,我们显式量化了收敛速率,并建立了足够光滑系数下的最优一阶收敛率。此外,我们证明了在高概率下对采样噪声一致成立的1/2阶弱收敛率,并在无波动率控制情况下建立了每条路径实现的1/2阶路径收敛性。基于这些结果,我们分析了基于离散观测的各种策略评估与策略梯度估计器的偏差与方差。我们的结果为 [H. Wang, T. Zariphopoulou, and X.Y. Zhou, J. Mach. Learn. Res., 21 (2020), pp. 1-34] 中探索性随机控制框架提供了理论支持。
原文摘要 · Abstract (English)
Stochastic policies (also known as relaxed controls) are widely used in continuous-time reinforcement learning algorithms. However, executing a stochastic policy and evaluating its performance in a continuous-time environment remain open challenges. This work introduces and rigorously analyzes a policy execution framework that samples actions from a stochastic policy at discrete time points and implements them as piecewise constant controls. We prove that as the sampling mesh size tends to zero, the controlled state process converges weakly to the dynamics with coefficients aggregated according to the stochastic policy. We explicitly quantify the convergence rate based on the regularity of the coefficients and establish an optimal first-order convergence rate for sufficiently regular coefficients. Additionally, we prove a $1/2$-order weak convergence rate that holds uniformly over the sampling noise with high probability, and establish a $1/2$-order pathwise convergence for each realization of the system noise in the absence of volatility control. Building on these results, we analyze the bias and variance of various policy evaluation and policy gradient estimators based on discrete-time observations. Our results provide theoretical justification for the exploratory stochastic control framework in [H. Wang, T. Zariphopoulou, and X.Y. Zhou, J. Mach. Learn. Res., 21 (2020), pp. 1-34].
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。