arXiv:2409.17200cs.LGmath.PR2024-09被引 1

用随机测度建模连续时间强化学习中的探索行为。

A random measure approach to reinforcement learning in continuous time

  • 将采样控制转化为由布朗运动和泊松测度驱动的随机测度形式。
  • 证明网格细化极限下,随机测度收敛到白噪声与泊松测度共同驱动的SDE。
  • 新模型可替代现有连续时间RL的探索与采样SDE,适用于理论分析与算法推导。

我们提出一种随机测度方法,用于建模连续时间强化学习中带有受控扩散和跳跃的探索行为,即执行测度值控制。首先,在连续时间中对随机控制进行离散网格采样,将由此产生的随机微分方程(SDE)重构为由适当随机测度驱动的方程。这些随机测度的构造利用了原始模型动力学中的布朗运动和泊松随机测度,以及在网格上采样的额外随机变量。接着,我们证明了当采样网格的粒度趋于零时,这些随机测度的极限定理,从而得到由白噪声随机测度和泊松随机测度共同驱动的网格采样极限SDE。我们还论证该极限SDE可替代近期连续时间强化学习文献中的探索性SDE和采样SDE,可用于探索性控制问题的理论分析及学习算法的推导。

原文摘要 · Abstract (English)

We present a random measure approach for modeling exploration, i.e., the execution of measure-valued controls, in continuous-time reinforcement learning (RL) with controlled diffusion and jumps. First, we consider the case when sampling the randomized control in continuous time takes place on a discrete-time grid and reformulate the resulting stochastic differential equation (SDE) as an equation driven by suitable random measures. The construction of these random measures makes use of the Brownian motion and the Poisson random measure (which are the sources of noise in the original model dynamics) as well as the additional random variables, which are sampled on the grid for the control execution. Then, we prove a limit theorem for these random measures as the mesh-size of the sampling grid goes to zero, which leads to the grid-sampling limit SDE that is jointly driven by white noise random measures and a Poisson random measure. We also argue that the grid-sampling limit SDE can substitute the exploratory SDE and the sample SDE of the recent continuous-time RL literature, i.e., it can be applied for the theoretical analysis of exploratory control problems and for the derivation of learning algorithms.

强化学习随机测度连续时间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。