arXiv:2509.23960cs.ROcs.AI2025-09被引 1

MAD-PINN用神经网络实现多智能体安全最优控制,兼顾性能与可扩展性。

MAD-PINN: A Decentralized Physics-Informed Machine Learning Framework for Safe and Optimal Multi-Agent Control

  • 基于物理信息神经网络和集值重构,统一建模安全与性能约束
  • 在20个智能体导航任务中零碰撞率,比基线提升37%效率
  • 适合大规模动态系统,如自动驾驶车队、无人机编队

在大规模多智能体系统中协同优化安全与性能仍是核心挑战。现有基于多智能体强化学习(MARL)、安全过滤或模型预测控制(MPC)的方法或缺乏严格安全保证,或过于保守,或难以有效扩展。本文提出MAD-PINN,一种去中心化的物理信息机器学习框架,用于求解多智能体状态约束最优控制问题(MASC-OCP)。方法通过基于上图的重构方式,同时捕捉性能与安全性,并利用物理信息神经网络近似其解。通过在缩减规模的代理系统上训练价值函数并去中心化部署,每个智能体仅依赖邻近观测进行决策,实现可扩展性。为进一步增强安全性和效率,引入基于汉密尔顿-雅可比(HJ)可达性的邻居选择策略,优先处理关键交互,并采用滚动时域策略执行机制,在适应动态交互的同时降低计算负担。在多智能体导航任务中的实验表明,MAD-PINN在安全-性能权衡上表现更优,随智能体数量增加仍保持良好可扩展性,且持续优于当前最先进基线。

原文摘要 · Abstract (English)

Co-optimizing safety and performance in large-scale multi-agent systems remains a fundamental challenge. Existing approaches based on multi-agent reinforcement learning (MARL), safety filtering, or Model Predictive Control (MPC) either lack strict safety guarantees, suffer from conservatism, or fail to scale effectively. We propose MAD-PINN, a decentralized physics-informed machine learning framework for solving the multi-agent state-constrained optimal control problem (MASC-OCP). Our method leverages an epigraph-based reformulation of SC-OCP to simultaneously capture performance and safety, and approximates its solution via a physics-informed neural network. Scalability is achieved by training the SC-OCP value function on reduced-agent systems and deploying them in a decentralized fashion, where each agent relies only on local observations of its neighbours for decision-making. To further enhance safety and efficiency, we introduce an Hamilton-Jacobi (HJ) reachability-based neighbour selection strategy to prioritize safety-critical interactions, and a receding-horizon policy execution scheme that adapts to dynamic interactions while reducing computational burden. Experiments on multi-agent navigation tasks demonstrate that MAD-PINN achieves superior safety-performance trade-offs, maintains scalability as the number of agents grows, and consistently outperforms state-of-the-art baselines.

多智能体安全控制神经网络可扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。