提出混合神经算子,实现带状态约束的多智能体博弈价值函数高效泛化逼近。
Parametric Value Approximation for General-sum Differential Games with State Constraints
- 设计混合神经算子,融合监督数据与物理驱动数据,映射游戏参数到价值函数。
- 在9维和13维非线性系统中,相同资源下安全性能优于传统监督算子。
- 适用于人机协作等需实时决策的安全关键场景,支持参数空间泛化。
一般和微分博弈可通过哈密顿-雅可比-伊斯阿克斯(HJI)方程近似值函数,实现在信息不完全下的高效推理。然而,传统方法面临维度灾难(CoD)问题。物理信息神经网络(PINNs)可缓解维度灾难并近似值函数,但在状态约束导致值函数具有大利普希茨常数时,其收敛性存在问题,尤其在安全关键应用中。此外,需在游戏参数空间上学习可泛化的值函数,而非为每种玩家类型单独训练。为此,本文提出混合神经算子(HNO),一种将博弈参数函数映射为值函数的算子。HNO利用有监督数据及全时空域的偏微分方程驱动数据进行模型精炼。在9维和13维非线性动态系统与状态约束下评估,对比基于DeepONet的监督神经算子(SNO),在相同计算预算与训练数据下,HNO在安全性能上表现更优。该工作推动了复杂人机或多智能体交互中可扩展、可泛化的值函数逼近,支持实时推理。
原文摘要 · Abstract (English)
General-sum differential games can approximate values solved by Hamilton-Jacobi-Isaacs (HJI) equations for efficient inference when information is incomplete. However, solving such games through conventional methods encounters the curse of dimensionality (CoD). Physics-informed neural networks (PINNs) offer a scalable approach to alleviate the CoD and approximate values, but there exist convergence issues for value approximations through vanilla PINNs when state constraints lead to values with large Lipschitz constants, particularly in safety-critical applications. In addition to addressing CoD, it is necessary to learn a generalizable value across a parametric space of games, rather than training multiple ones for each specific player-type configuration. To overcome these challenges, we propose a Hybrid Neural Operator (HNO), which is an operator that can map parameter functions for games to value functions. HNO leverages informative supervised data and samples PDE-driven data across entire spatial-temporal space for model refinement. We evaluate HNO on 9D and 13D scenarios with nonlinear dynamics and state constraints, comparing it against a Supervised Neural Operator (a variant of DeepONet). Under the same computational budget and training data, HNO outperforms SNO for safety performance. This work provides a step toward scalable and generalizable value function approximation, enabling real-time inference for complex human-robot or multi-agent interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。