arXiv:2504.08195cs.MAcs.AI2025-04中稿 · the 2025 IEEE Inte…被引 9

用图神经网络与注意力机制提升无人机群协同任务效率

Graph Based Deep Reinforcement Learning Aided by Transformers for Multi-Agent Cooperation

  • 构建动态图模型捕捉无人机与目标间关系,融合Transformer消息传递
  • 90%任务完成率,100%区域覆盖,每轮平均步数降至200(基准600)
  • 适合复杂环境下多智能体协同规划,尤其通信受限场景

在灾难救援、环境监测等需服务分散目标点的应用中,大量自主无人机协同执行任务面临部分可观测性、通信范围有限及环境不确定等挑战。传统路径规划算法在缺乏先验信息时表现不佳。为此,我们提出一种融合图神经网络(GNN)、深度强化学习(DRL)与基于Transformer的消息传递机制的新框架,以增强多智能体协作能力。该方法通过自适应图结构建模智能体-智能体及智能体-目标间的交互,实现信息高效聚合与决策;引入边缘特征增强的注意力机制,捕捉复杂交互模式;结合双深度Q网络(Double DQN)与优先经验回放,优化部分可观测环境下的策略。实验表明,该方法在服务覆盖率90%、网格覆盖率达100%(节点发现率)的同时,将每轮平均步数降至200,显著优于粒子群优化(PSO)、贪心算法及DQN等基准方法。

原文摘要 · Abstract (English)

Mission planning for a fleet of cooperative autonomous drones in applications that involve serving distributed target points, such as disaster response, environmental monitoring, and surveillance, is challenging, especially under partial observability, limited communication range, and uncertain environments. Traditional path-planning algorithms struggle in these scenarios, particularly when prior information is not available. To address these challenges, we propose a novel framework that integrates Graph Neural Networks (GNNs), Deep Reinforcement Learning (DRL), and transformer-based mechanisms for enhanced multi-agent coordination and collective task execution. Our approach leverages GNNs to model agent-agent and agent-goal interactions through adaptive graph construction, enabling efficient information aggregation and decision-making under constrained communication. A transformer-based message-passing mechanism, augmented with edge-feature-enhanced attention, captures complex interaction patterns, while a Double Deep Q-Network (Double DQN) with prioritized experience replay optimizes agent policies in partially observable environments. This integration is carefully designed to address specific requirements of multi-agent navigation, such as scalability, adaptability, and efficient task execution. Experimental results demonstrate superior performance, with 90% service provisioning and 100% grid coverage (node discovery), while reducing the average steps per episode to 200, compared to 600 for benchmark methods such as particle swarm optimization (PSO), greedy algorithms and DQN.

多智能体强化学习图神经网络无人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。