提出连续时间强化学习方法,解决非马尔可夫型跳跃扩散过程的控制问题。
Continuous-Time Reinforcement Learning for Controlled Hawkes Jump-Diffusions

- 通过指数混合近似将多变量霍克斯过程马尔可夫化
- 在三种核函数下性能优于离散时间强化学习方法
- 仅需事件时间与路径数据,无需已知核系数
我们研究在非马尔可夫设置下,基于机器学习算法对多变量霍克斯驱动随机微分方程进行随机控制。由于霍克斯强度的记忆路径依赖性,该问题不适用于经典随机控制理论,除非特定马尔可夫核。我们首先开发了一种有限维马尔可夫化程序和算法,用指数核混合逼近多变量霍克斯过程,并证明了该近似在过程、强度及问题价值上均收敛到原始非马尔可夫过程及其原问题价值。随后,我们在该马尔可夫化近似基础上构建连续时间确定性策略梯度学习方法,称为霍克斯-连续时间DDPG(Hawkes-CT DDPG)。我们提出一种无需模型的算法,仅通过观测过程的事件时间、SDE解的实现以及一组选定衰减滤波器,即可求解非马尔可夫霍克斯驱动优化问题,而无需知晓霍克斯核系数。我们在三种不同核函数(简单指数、爱尔朗、幂律)下,将我们的连续时间强化学习方法与离散时间强化学习技术进行了比较。
原文摘要 · Abstract (English)
We study stochastic control of multivariate Hawkes-driven stochastic differential equations with machine learning algorithms in a non-Markovian setting. Due to the path dependence of the memory of the Hawkes intensity, this problem does not fall within classical stochastic control theory outside particular Markovian kernels. We first develop a finite-dimensional Markovianization procedure and algorithm to approximate multivariate Hawkes processes with mixtures of exponential kernels. We prove the convergence of the Markovianized approximation of the Hawkes process, its intensity, and the value of the problem to the original non-Markovian processes and the value of the primal problem. We then formulate continuous-time deterministic policy gradient learning on the Markovianized approximation of the problem, called Hawkes-CT DDPG. We propose a model-free algorithm to solve the non-Markovian Hawkes-driven optimization by observing only the event times of the process, the realization of the solution to the SDE, and a chosen set of decay filters, while the Hawkes kernel coefficients remain unknown. We compare our continuous time reinforcement learning Hawkes-CT DDPG method with discrete time reinforcement learning techniques under three different types of kernels: simple exponential, Erlang, and power-law kernels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。