arXiv:2606.28671cs.LG2026-06被引 1

用熵正则化强化学习解决随机切换下的层级博弈问题

Entropy-Regularized Reinforcement Learning for Linear-Quadratic Stackelberg Differential Games in Regime-Switching Diffusion Models

论文配图:Entropy-Regularized Reinforcement Learning for Linear-Quadratic Stackelberg Differential Games in Regime-Switching Diffusion Models
图 1 · 摘自论文原文
  • 引入熵正则化构造可探索的弱耦合哈密顿方程,避免陷入次优解
  • 神经网络逼近分段值函数,在高维下求解效率显著提升
  • 适合研究带突变环境的多智能体决策系统

Stackelberg微分博弈(SDGs)为随机连续时间环境中的层级决策提供了强大框架,但传统动态规划与哈密顿-雅可比-贝尔曼-伊斯阿克斯(HJBI)方法在高维系统中计算复杂度高,求解困难。本文提出一种熵正则化强化学习(ERRL)方法,用于马尔可夫制度切换扩散模型下的线性二次型堆叠贝格微分博弈(LQ-SDGs)。核心创新在于推导出带有熵正则化的探索性弱耦合HJBI方程,促进随机策略主动避开次优均衡——这是经典方法的缺陷。通过神经网络逼近分段依赖制度的值函数,高效求解高维偏微分方程(PDE),并引入新型采样技术提升计算可行性。数值结果表明,该框架相比传统方法更有效,尤其在通过探索性策略摆脱次优陷阱方面表现突出。研究强调了熵正则化与神经网络近似在应对突发环境变化的层级决策问题中的关键作用。

原文摘要 · Abstract (English)

Stackelberg differential games (SDGs) provide a powerful framework for hierarchical decision-making in stochastic and continuous-time environments, yet their solution remains computationally challenging due to the complexity of traditional dynamic programming and Hamilton-Jacobi-Bellman-Isaacs (HJBI) methods, especially in high-dimensional systems. This paper proposes an entropy-regularized reinforcement learning (ERRL) approach for linear-quadratic SDGs (LQ-SDGs) within a continuous-time diffusion framework governed by Markovian regime switching. The key innovation lies in deriving exploratory weakly-coupled HJBI equations with entropy regularization, which promotes stochastic policies that actively avoid suboptimal equilibria -- a limitation of classical SDG methods. Neural networks are integrated to approximate regime-dependent value functions and solve high-dimensional partial differential equations (PDEs) efficiently, while a novel sampling technique enhances computational tractability. Numerical results demonstrate the effectiveness of the framework compared to conventional approaches, particularly in escaping suboptimal traps through exploratory policies. The study highlights the critical role of entropy regularization and neural network approximations in achieving robust solutions for hierarchical decision-making problems under abrupt environmental shifts.

强化学习博弈论随机控制神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。