分布式多智能体系统中安全高效参数调优的新方法
Towards safe control parameter tuning in distributed multi-agent systems
- 用带时空核的贝叶斯优化处理非凸约束下的分布式优化
- 通过邻近通信建模全局行为,实现样本高效收敛
- 适合自动驾驶、协作机器人等安全关键场景
许多现实世界的安全关键问题,如自动驾驶和协作机器人,具有分布式多智能体特性。为在保证安全的前提下优化系统性能,可将问题建模为分布式优化:每个智能体优化自身参数以最大化耦合奖励函数,并满足耦合约束。现有工作或采用集中式设置,或忽略安全性,或存在采样效率低的问题。由于需兼顾采样效率且面对未知的非凸奖励与约束,本文采用基于高斯过程回归的安全贝叶斯优化。同时考虑智能体间的最近邻通信,为捕捉非邻近智能体的行为,将静态全局优化问题重构为每个智能体的时变局部优化问题,引入时间作为隐变量。为此,提出一种定制的时空核以融合先验知识。仿真结果验证了算法的有效性。
原文摘要 · Abstract (English)
Many safety-critical real-world problems, such as autonomous driving and collaborative robots, are of a distributed multi-agent nature. To optimize the performance of these systems while ensuring safety, we can cast them as distributed optimization problems, where each agent aims to optimize their parameters to maximize a coupled reward function subject to coupled constraints. Prior work either studies a centralized setting, does not consider safety, or struggles with sample efficiency. Since we require sample efficiency and work with unknown and nonconvex rewards and constraints, we solve this optimization problem using safe Bayesian optimization with Gaussian process regression. Moreover, we consider nearest-neighbor communication between the agents. To capture the behavior of non-neighboring agents, we reformulate the static global optimization problem as a time-varying local optimization problem for each agent, essentially introducing time as a latent variable. To this end, we propose a custom spatio-temporal kernel to integrate prior knowledge. We show the successful deployment of our algorithm in simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。