提出TRIDENT框架,解决多智能体强化学习中的安全与物理耦合难题。
TRIDENT: Breaking the Hybrid-Safety-Physics Coupling for Provably Safe Multi-Agent Reinforcement Learning

- 三组件协同设计,破解离散连续动作与安全约束的耦合问题。
- 训练违规减少95.5%,奖励提升13.5%,收敛至受限纳什均衡。
- 适合需高安全性与物理一致性场景,如无人机群、交通管控。
网络化人机系统中的安全协作要求学习算法同时处理混合离散-连续动作、严格的训练期安全约束及物理驱动的动力学。我们揭示这三者形成定向偏差循环,阻碍现有模块简单组合,并形式化为三重耦合引理。为此提出TRIDENT,首个三部分协同设计的多智能体强化学习框架:基于Richardson-Romberg梯度修正将Gumbel-Softmax偏差从O(tau)降至O(tau²),采用李雅普诺夫约束的序列信任域更新确保每步可行性,以及物理信息残差价值函数分解值函数而非奖励。证明其收敛率O~(1/sqrt(K))至受限纳什均衡,累积违规量为O(sqrt(K))。在多无人机移动边缘计算、自主交叉路口管理及混合SMAC变体上,相比MADDPG减少95.5%训练违规,比MACPO减少76.3%,同时奖励高于最强无约束基线13.5%。
原文摘要 · Abstract (English)
Safe coordination in networked cyber-physical systems forces learning algorithms to simultaneously handle hybrid discrete-continuous actions, hard training-time safety constraints, and physics-governed dynamics. We show that these three features form a directed cycle of biases that defeats any naive composition of off-the-shelf modules, and formalize this as a three-way coupling lemma. We then introduce TRIDENT, the first MARL framework whose three components are co-designed to cancel each leak: a Richardson-Romberg gradient correction reducing Gumbel-Softmax bias from O(tau) to O(tau^2), a Lyapunov-constrained sequential trust-region update enforcing per-iterate feasibility, and a physics-informed residual critic that decomposes value rather than reward. We prove an O~(1/sqrt(K)) convergence rate to a constrained Nash equilibrium and an O(sqrt(K)) cumulative-violation bound. On multi-UAV mobile-edge computing, autonomous intersection management, and a hybrid SMAC variant, TRIDENT cuts training-time violations by 95.5% over MADDPG and 76.3% over MACPO, while improving reward by 13.5% over the strongest unconstrained baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。