arXiv:2505.00909cs.LGmath.OC2025-05被引 2

用高斯过程与叠加萨瓦尔加速,高效求解哈密顿-雅可比-贝尔曼方程和平均场博弈的正反问题。

Gaussian process policy iteration with additive Schwarz acceleration for forward and inverse HJB and mean field game problems

  • 基于高斯过程建模未知场,将非线性问题转为一系列线性偏微分方程子问题。
  • 通过勒让德变换点对点更新策略,支持二次、凸及带约束控制成本的优化。
  • 引入叠加萨瓦尔预处理加速收敛,显著提升计算效率,适用于复杂动态系统建模。

本文提出一种基于高斯过程(GP)的策略迭代框架,用于求解哈密顿-雅可比-贝尔曼(HJB)方程和平均场博弈(MFG)中的正向与逆向问题。策略迭代被构造成在固定策略下评估价值函数与改进策略的交替过程。在该框架中,利用高斯过程对未知场进行建模,将非线性系统转化为一系列线性偏微分方程(PDE)子问题。借助线性结构,在线性PDE插值约束下,价值函数与(在MFG设定中)群体密度的更新具有显式表示公式。策略随后通过勒让德变换步骤点对点更新,涉及对控制变量的低维最大化。对于标准二次代价函数,该最大化显式可解;对于光滑严格凸代价,通过一阶最优性条件求解;在约束或非光滑情况下,则退化为低维约束最大化问题。为提升收敛速度,每轮策略更新后引入叠加萨瓦尔(additive Schwarz)加速作为预处理步骤。数值实验验证了该加速方法在提升计算效率方面的有效性。

原文摘要 · Abstract (English)

In this paper, we propose a Gaussian Process (GP)-based policy iteration framework for addressing both forward and inverse problems in Hamilton--Jacobi--Bellman (HJB) equations and mean field games (MFGs). Policy iteration is formulated as an alternating procedure between evaluating the value function under a fixed control policy and improving the policy. In our approach, we model the unknown fields using GPs within a policy-iteration framework that converts the nonlinear system into a sequence of linear PDE subproblems. Then, leveraging the linear structure, the updates for the value function and, in the MFG setting, the population density admit explicit representer formulas under linear PDE collocation constraints. The policy is subsequently updated pointwise via a Legendre transform step, which involves a low-dimensional maximization over the control variable. This maximization is explicit for standard quadratic costs. For smooth, strictly convex costs, this pointwise maximization is solved through its first-order optimality condition, whereas in constrained or non-smooth cases, it becomes a low-dimensional constrained maximization problem. To improve convergence, we incorporate the additive Schwarz acceleration as a preconditioning step following each policy update. Numerical experiments demonstrate the effectiveness of the Schwarz acceleration in improving computational efficiency.

高斯过程策略迭代平均场博弈偏微分方程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。