用归一化流构建连续时间跳跃控制的强化学习框架,高效求解金融优化难题。
An Actor-Critic Framework for Continuous-Time Jump-Diffusion Controls with Normalizing Flows
- 基于时变小q函数与占位测度,构建可处理跳跃与时变参数的策略梯度方法
- 在多资产投资组合等任务中实现稳定学习与高维问题的良好扩展性
- 适合金融工程、量化交易等领域研究者,尤其关注非高斯策略建模场景
具有时变跳跃扩散动态的连续时间随机控制在金融与经济中至关重要,但显式时间依赖、不连续冲击和高维性使得最优策略计算困难。本文提出一种无需网格的演员-评论家框架,用于求解熵正则化控制问题与含跳跃的随机博弈。方法基于时变小q函数与合适的占位测度,获得可容纳时变漂移、波动率与跳跃项的策略梯度表示。为在连续动作空间中表达灵活的随机策略,演员采用条件归一化流参数化,既能实现非高斯策略,又保持精确似然评估以支持熵正则化与策略优化。在时变线性二次控制、默顿投资组合优化及多智能体投资组合博弈中验证方法,使用解析解或高精度基准进行对比。数值结果表明,在跳跃不连续情况下学习稳定,能准确逼近最优随机策略,并在维度与智能体数量上具有良好扩展性。
原文摘要 · Abstract (English)
Continuous-time stochastic control with time-inhomogeneous jump-diffusion dynamics is central in finance and economics, but computing optimal policies is difficult under explicit time dependence, discontinuous shocks, and high dimensionality. We propose an actor-critic framework that serves as a mesh-free solver for entropy-regularized control problems and stochastic games with jumps. The approach is built on a time-inhomogeneous little q-function and an appropriate occupation measure, yielding a policy-gradient representation that accommodates time-dependent drift, volatility, and jump terms. To represent expressive stochastic policies in continuous-action spaces, we parameterize the actor using conditional normalizing flows, enabling flexible non-Gaussian policies while retaining exact likelihood evaluation for entropy regularization and policy optimization. We validate the method on time-inhomogeneous linear-quadratic control, Merton portfolio optimization, and a multi-agent portfolio game, using explicit solutions or high-accuracy benchmarks. Numerical results demonstrate stable learning under jump discontinuities, accurate approximation of optimal stochastic policies, and favorable scaling with respect to dimension and number of agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。