融合强化学习与模型预测控制,提升多智能体协作的可靠性与安全性。
Merging model-based control with multi-agent reinforcement learning for multi-agent cooperative teaming strategies

- 用基于模型的预测控制增强多智能体强化学习策略的动态可行性。
- 硬件实验显示,新方法在协同降落任务中成功率达100%,优于60%的基线。
- 适合需要高安全性和实时响应的多机器人协作场景。
本文提出一种融合多智能体强化学习(MARL)与模型预测控制(MPC)的框架,实现多智能体协作任务中的安全、动态可行动作。多智能体强化学习可在长周期规划中从离散不可导奖励中学习协作策略;而模型预测控制则在短时重规划中提供鲁棒且安全的动作。我们提出一种扩展的演员-评论家模型预测控制算法,称为多智能体演员-评论家模型预测控制(MA-AC-MPC)。通过多智能体追逃场景验证该算法:对比使用MA-AC-MPC与多层感知机(MA-AC-MLP)的逃逸团队策略,追捕方采用增强比例导航法作为先进对抗控制律。在异构环境中,无人机与全向轮机器人协同完成着陆任务,硬件测试表明MA-AC-MPC成功率达100%,而MA-AC-MLP仅为60%。实验证明了该算法在两类环境中的鲁棒性与实用性。
原文摘要 · Abstract (English)
In this work, we propose a framework that combines multi-agent reinforcement learning (MARL) with model-based control to achieve safe, dynamically feasible actions in cooperative multi-agent tasks. Multi-agent reinforcement learning provides the advantage of learning cooperative policies for multi-agent teams from discrete non-differentiable rewards in a long planning horizon. Model-predictive control is robust and offers safe, dynamically feasible actions in a fast replanning framework for short horizons. We propose an algorithm that extends actor-critic model predictive control for MARL which we refer to as multi-agent actor-critic model predictive control (MA-AC-MPC). We demonstrate the capabilities of this algorithm by applying it to a multi-agent pursuit-evasion scenario. Specifically, we compare the evader team's strategy using the MA-AC-MPC model and a multi-layer perceptron model (MA-AC-MLP). The pursuer team uses augmented proportional navigation as it is accepted as an advanced adversarial control law. We also provide an example with a heterogeneous environment where a drone and omni-wheeled rover cooperate to achieve repeatable and successful landing with 100% success rate in hardware for MA-AC-MPC compared to 60% for MA-AC-MLP. We demonstrate the robustness of the proposed MA-AC-MPC algorithm in hardware for both environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。