用熵正则化强化学习解决环境突变下的博弈策略问题
Entropy Regularized Reinforcement Learning for Zero-Sum Stochastic Differential Games in a Regime-Switching Jump-Diffusion Process

- 将策略建模为状态相关的概率分布,增强鲁棒性
- 线性二次情形下可求解耦合常微分方程得解析解
- 适用于金融投资等存在制度切换的动态博弈场景
为应对传统随机微分博弈框架中的参数误设和突发结构性环境变化,本文提出一种分布式控制方法,将最优策略表示为在连续状态、离散制度状态及参数条件下的动作概率分布,构建了制度切换跳跃扩散过程下的熵正则化零和随机微分博弈(ERRL-ZSSDG)强化学习框架。基于动态规划原理,推导出关联的耦合哈密顿-雅可比-贝尔曼-伊斯阿克斯(HJBI)方程组,均衡策略由价值函数梯度表达。在线性二次情形下,通过求解耦合常微分方程组获得价值函数与均衡策略的半解析解。在更一般情形中,设计了演员-评论家策略改进算法,用于跨制度近似价值函数与均衡策略。方法应用于投资博弈,数值实验展示了温度参数与制度转换对最优策略与价值的影响。
原文摘要 · Abstract (English)
To address parameter misspecification and sudden structural environmental changes in conventional stochastic differential game (SDG) frameworks, this paper introduces a distributional control approach that characterizes optimal strategies as probability distributions over actions, conditioned on the continuous state, the discrete regime state, and parameters. This forms a reinforcement learning framework for entropy-regularized zero-sum stochastic differential games (ERRL-ZSSDGs) in a regime-switching jump-diffusion process. Using the dynamic programming principle (DPP), we derive the associated coupled systems of Hamilton-Jacobi-Bellman-Isaacs (HJBI) equations, from which equilibrium strategies are expressed via gradients of the value function. For linear-quadratic problems, semi-analytical solutions for both value function and equilibrium strategies are obtained by solving a system of coupled ordinary differential equations (ODEs). In more general settings, an Actor-Critic policy improvement algorithm is developed to approximate the value functions and equilibrium policies across different regimes. The method is applied to an investment game, and numerical examples illustrate the effect of the temperature parameter and regime transitions on optimal policies and values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。