arXiv:2603.16141cs.MAcs.LG2026-03

让无人机群在通信受限下仍能高效协作部署,提升覆盖范围。

Communication-Aware Multi-Agent Reinforcement Learning for Decentralized Cooperative UAV Deployment

  • 基于图注意力机制,融合局部状态与邻近节点消息进行决策。
  • 通信受限时相比MAPPO提升约12%目标覆盖率,接近全局最优解。
  • 适合需要低通信开销的分布式无人机协同任务场景。

自主无人机集群日益作为快速部署的空中中继和感知平台使用,但实际部署需在部分可观测性和间歇性点对点连接条件下运行。本文提出一种基于图的多智能体强化学习框架,采用集中训练、分散执行(CTDE)模式:训练时使用全局状态和集中式评价器,执行时每个无人机仅依赖本地观测和邻近节点消息。在受限通信下,邻居关系由信噪比阈值连接图定义。架构通过智能体-实体注意力模块编码本地状态与邻近实体信息,并利用基于信道模型的信号质量限制通信图,对无人机间消息进行邻居自注意力聚合。我们在合作中继部署任务DroneConnect上评估该框架,实验表明,在通信受限和部分可观测条件下,该方法相比MAPPO实现约12%的目标覆盖率提升,且在全节点可观测的混合整数线性规划(MILP)离线上界下仍具竞争力。

原文摘要 · Abstract (English)

Autonomous Unmanned Aerial Vehicle (UAV) swarms are increasingly used as rapidly deployable aerial relays and sensing platforms, yet practical deployments must operate under partial observability and intermittent peer-to-peer connectivity. We present a graph-based multi-agent reinforcement learning framework trained under centralized training with decentralized execution (CTDE): a centralized critic and global state are available only during training, while each UAV executes a shared policy using local observations and messages from nearby neighbors. Under restricted communication, neighbor relations are induced by an SNR-threshold connectivity graph. Our architecture encodes local agent state and nearby entities with an agent-entity attention module and aggregates inter-UAV messages with neighbor self-attention over a signal-quality-limited communication graph defined by a channel model. We evaluate the framework on a cooperative relay-deployment task, DroneConnect. Experimental results show that the proposed method achieves an approximately 12% increase in target coverage over MAPPO under restricted communication and partial observability, while remaining competitive with a mixed-integer linear programming (MILP)-based offline upper bound with full node observability.

多智能体无人机群强化学习通信受限

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。