arXiv:2606.02107cs.ROcs.AI2026-06中稿 · Manuscript version…

让无人机群在通信受限下自动对齐,无需重训就能扩展到250架。

Network Distributed Multi-Agent Reinforcement Learning for Consensus Control of Quadcopters

论文配图:Network Distributed Multi-Agent Reinforcement Learning for Consensus Control of Quadcopters
图 1 · 摘自论文原文
  • 将通信拓扑融入决策,每架无人机只看两个邻居信息。
  • 三机训练的策略可直接用于250架无人机,收敛稳定。
  • 适合大规模、低通信量的无人机协同控制场景。

本文提出一种网络分布式多智能体强化学习(ND-MARL)框架,用于四旋翼无人机群的一致性控制。与依赖中心化规划或完全去中心化执行的传统多智能体强化学习不同,ND-MARL将群体通信图整合进决策过程。在2-邻居通信拓扑下,每个智能体仅观测两个邻近智能体信息,并通过分布式策略输出动作。高层分布式一致性规划器使用多智能体软演员-评论家(MASAC)训练,嵌入分层结构中,生成参考目标位置供低层四旋翼控制器跟踪。实验表明,相比集中式MARL控制器,该方法实现了平滑的一致性轨迹及良好的规划-跟踪融合效果。最显著的是,所学控制器具备零样本可扩展性:在三智能体系统上训练的策略可直接部署至最多250个智能体的群体中,无需重新训练或微调,在相同2-邻居通信拓扑下仍能实现一致收敛;随着团队规模增大,稳态散布略有增加,源于信息传播稀疏性。这些结果表明,ND-MARL是一种稳定、通信感知的分布式四旋翼一致性控制框架。

原文摘要 · Abstract (English)

This paper proposes a Network Distributed Multi-Agent Reinforcement Learning (ND-MARL) framework for quadcopter consensus control. Compared to conventional multi-agent MARL formulations that rely on centralized planning or fully decentralized execution, ND-MARL incorporates the swarm communication graph into the decision process. Under a 2-Neighbor communication topology, each agent observes information of only two neighbors and outputs an action through a distributed policy. A high-level distributed consensus planner is trained using Multi-Agent Soft Actor-Critic (MASAC) and embedded in a hierarchical stack to generate reference target positions tracked by a low-level quadcopter controller. Results demonstrate smooth consensus trajectories and planner-tracker integration when compared to a centralized MARL controller. Most notably, the learned controller exhibits zero-shot scalability, as policies trained on a three-agent system are deployed to swarms of up to 250 agents under the same 2-Neighbor communication topology without retraining or fine-tuning, achieving consistent convergence with increasing steady-state spread at large team sizes due to sparse information propagation. These findings highlight ND-MARL as a stable framework for distributed, communication-aware quadcopter consensus control.

多智能体无人机群强化学习分布式控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。