arXiv:2605.24992cs.NIcs.AI2026-05被引 3

为无人机群设计节能多智能体强化学习,提升任务成功率与效率。

Scaling up Energy-Aware Multi-Agent Reinforcement Learning for Mission-Oriented Drone Networks with Individual Reward

论文配图:Scaling up Energy-Aware Multi-Agent Reinforcement Learning for Mission-Oriented Drone Networks with Individual Reward
图 1 · 摘自论文原文
  • 采用个体奖励机制,结合任务进度与剩余电量优化决策。
  • 在高任务密度下成功率达近100%,环境规模扩大时表现更稳定。
  • 适合大规模动态无人机网络,尤其关注能效与鲁棒性场景。

多智能体强化学习(MARL)在自动驾驶、智慧城市等协同系统中展现出广泛应用潜力,近年来也被用于解决无人机网络的轨迹规划问题。然而,动态环境与电池容量有限仍是实现高效协作任务执行的挑战。本文提出一种能量感知的MARL模型,基于深度Q网络(DQN)并引入由任务进展和剩余电量驱动的个体奖励函数。通过模拟实验对比共享奖励的MARL方法,验证了信用分配策略的影响。结果表明,所提模型在任意任务位置与长度下均能达到至少80%的成功率;当任务密度接近40时,成功率可逼近100%。其核心优势在于环境扩展时表现更稳健,相比共享奖励模型,在更大规模场景中以更少步数达成更高成功率,目标明确性提升了能源效率。

原文摘要 · Abstract (English)

Multi-agent reinforcement learning (MARL) has shown wide applicability in collaborative systems such as autonomous driving and smart cities for its ability of learning through interaction. With the recent development of drone networks, researchers have also applied MARL to address the trajectory planning problems. However, the dynamic environment and the limited battery capacity are still challenging for using MARL to achieve efficient collaborative task execution. In this paper, we propose an energy-aware MARL model as an attempt to tackle these challenges, leveraging Deep Q-Networks (DQN) with \emph{individual reward functions} driven by the task execution progress and the remaining battery of drones. We conduct a set of simulation studies for the proposed mode and compare it with the shared reward MARL~\cite{Li2022MARL} to explore the impact of credit assignment in MARL. The results indicate that our proposed model can achieve at least 80\% success rate regardless of the task locations and lengths. Similar to the shared reward mode, the individual reward mode can achieve a better success rate when the task density is high, and it can hit nearly a 100\% success rate when task density gets close to 40\%. The true advantage of our proposed model with individual reward is revealed when scaling up the environment. The comparison to the shared reward MARL shows that the our proposed model is more robust towards the change of the environment size and agent numbers. It can achieve higher success rate with fewer steps due to the clarity of the goal which improves energy efficiency even better.

无人机强化学习能源效率多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。