用深度学习求解带跳跃的连续时间随机控制问题,精度高且可扩展。
Deep Learning for Continuous-Time Stochastic Control with Jumps
- 用两个神经网络分别拟合最优策略和价值函数,迭代训练。
- 基于哈密顿-雅可比-贝尔曼方程设计双目标训练机制。
- 适用于高维复杂随机控制任务,尤其适合有跳跃过程的系统建模。
本文提出一种基于模型的深度学习方法,用于求解有限时域的连续时间随机控制问题(含跳跃)。通过迭代训练两个神经网络:一个表示最优策略,另一个逼近价值函数。利用连续时间动态规划原理,从哈密顿-雅可比-贝尔曼方程导出两种不同的训练目标,确保网络能捕捉底层随机动态。在多个问题上的实证评估表明,该方法具有高精度和良好的可扩展性,有效解决了高维复杂随机控制任务。
原文摘要 · Abstract (English)
In this paper, we introduce a model-based deep-learning approach to solve finite-horizon continuous-time stochastic control problems with jumps. We iteratively train two neural networks: one to represent the optimal policy and the other to approximate the value function. Leveraging a continuous-time version of the dynamic programming principle, we derive two different training objectives based on the Hamilton-Jacobi-Bellman equation, ensuring that the networks capture the underlying stochastic dynamics. Empirical evaluations on different problems illustrate the accuracy and scalability of our approach, demonstrating its effectiveness in solving complex high-dimensional stochastic control tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。