用深度学习解决连续时间委托代理问题,支持多维状态与约束。
DeepPAAC: A New Deep Galerkin Method for Principal-Agent Problems
- 提出DeepPAAC算法,基于深度强化学习求解带隐式哈密顿量的方程。
- 在五个案例中验证可处理多维状态与控制变量,收敛性良好。
- 适合金融工程、智能合约等需建模复杂激励机制的研究者。
我们研究连续时间委托代理(PA)问题的数值求解方法。构建了一个包含连续支付与一次性支付、且代理人策略为多维的通用PA模型。针对由此产生的具有隐式哈密顿量的哈密顿-雅可比-贝尔曼方程,提出一种新型深度学习方法:深度委托代理演员-评论家(DeepPAAC)算法。该方法能有效处理多维状态与控制变量,以及各类约束条件。通过五个不同案例研究,探讨了神经网络架构、训练设计、损失函数等对求解器收敛性的影响。
原文摘要 · Abstract (English)
We consider numerical resolution of principal-agent (PA) problems in continuous time. We formulate a generic PA model with continuous and lump payments and a multi-dimensional strategy of the agent. To tackle the resulting Hamilton-Jacobi-Bellman equation with an implicit Hamiltonian we develop a novel deep learning method: the Deep Principal-Agent Actor Critic (DeepPAAC) Actor-Critic algorithm. DeepPAAC is able to handle multi-dimensional states and controls, as well as constraints. We investigate the role of the neural network architecture, training designs, loss functions, etc. on the convergence of the solver, presenting five different case studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。