提出DGPPO框架,实现多智能体系统在未知动态下的安全高效控制。
Discrete GCBF Proximal Policy Optimization for Multi-agent Safe Optimal Control
- 基于离散图CBF与近端策略优化联合学习安全策略
- 在多种环境中实现高任务性能与高安全率,超参恒定
- 适用于动态拓扑、输入约束等复杂场景,适合工业级多智能体应用
能够实现高性能并满足安全约束的控制策略对任何系统(包括多智能体系统)都至关重要。分布式控制屏障函数(CBF)是一种有前景的安全保障方法。然而,在未知离散时间动力学、部分可观测性、邻域变化和输入约束条件下,设计基于分布式CBF的策略仍具挑战性,尤其当缺乏高性能的名义策略时。为此,本文提出DGPPO框架,同时学习一个能处理邻域变化和输入约束的离散图CBF,以及一个针对未知离散时间动力学的分布式高性能安全策略。我们在三个不同仿真引擎的多智能体任务上进行了实证验证。结果表明,相比现有方法,本框架所获得的策略在任务性能上达到基准水平(忽略安全约束的模型),在安全率上匹配最保守基准,且所有环境均使用固定超参数。
原文摘要 · Abstract (English)
Control policies that can achieve high task performance and satisfy safety constraints are desirable for any system, including multi-agent systems (MAS). One promising technique for ensuring the safety of MAS is distributed control barrier functions (CBF). However, it is difficult to design distributed CBF-based policies for MAS that can tackle unknown discrete-time dynamics, partial observability, changing neighborhoods, and input constraints, especially when a distributed high-performance nominal policy that can achieve the task is unavailable. To tackle these challenges, we propose DGPPO, a new framework that simultaneously learns both a discrete graph CBF which handles neighborhood changes and input constraints, and a distributed high-performance safe policy for MAS with unknown discrete-time dynamics. We empirically validate our claims on a suite of multi-agent tasks spanning three different simulation engines. The results suggest that, compared with existing methods, our DGPPO framework obtains policies that achieve high task performance (matching baselines that ignore the safety constraints), and high safety rates (matching the most conservative baselines), with a constant set of hyperparameters across all environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。