arXiv:2510.03823cs.LGcs.MA2025-10

用强化学习让高空气球自动分区域覆盖,效果媲美最优几何算法。

Distributed Area Coverage with High Altitude Balloons Using Multi-Agent Reinforcement Learning

  • 基于QMIX的多智能体强化学习,实现气球协同覆盖。
  • 在小团队任务中性能接近理论最优解,覆盖效率高。
  • 适合复杂自主任务,替代传统难解的确定性算法。

高空气球(HAB)可利用平流层风层实现有限水平控制,适用于侦察、环境监测和通信网络。现有协作方法多采用确定性策略(如Voronoi划分、极值搜索控制),在小型团队和局部任务中表现不佳。尽管单个气球的强化学习控制已有验证,但多智能体强化学习(MARL)在该领域尚未研究。本文首次系统应用多智能体强化学习(MARL)于HAB分布式区域覆盖任务。我们扩展了先前的强化学习仿真环境RLHAB,支持多智能体协同学习,使多个代理在真实大气条件下同时运行。采用QMIX算法,结合集中训练、分散执行机制,应对气球协调挑战。设计包含个体状态、环境上下文与队友信息的观测空间,采用分层奖励函数优先保证覆盖范围并鼓励空间分布。实验表明,QMIX在分布式区域覆盖任务中表现接近理论最优的几何确定性方法,验证了MARL的有效性,为更复杂的自主多气球任务提供了基础,尤其适用于传统确定性方法难以求解的场景。

原文摘要 · Abstract (English)

High Altitude Balloons (HABs) can leverage stratospheric wind layers for limited horizontal control, enabling applications in reconnaissance, environmental monitoring, and communications networks. Existing multi-agent HAB coordination approaches use deterministic methods like Voronoi partitioning and extremum seeking control for large global constellations, which perform poorly for smaller teams and localized missions. While single-agent HAB control using reinforcement learning has been demonstrated on HABs, coordinated multi-agent reinforcement learning (MARL) has not yet been investigated. This work presents the first systematic application of multi-agent reinforcement learning (MARL) to HAB coordination for distributed area coverage. We extend our previously developed reinforcement learning simulation environment (RLHAB) to support cooperative multi-agent learning, enabling multiple agents to operate simultaneously in realistic atmospheric conditions. We adapt QMIX for HAB area coverage coordination, leveraging Centralized Training with Decentralized Execution to address atmospheric vehicle coordination challenges. Our approach employs specialized observation spaces providing individual state, environmental context, and teammate data, with hierarchical rewards prioritizing coverage while encouraging spatial distribution. We demonstrate that QMIX achieves similar performance to the theoretically optimal geometric deterministic method for distributed area coverage, validating the MARL approach and providing a foundation for more complex autonomous multi-HAB missions where deterministic methods become intractable.

强化学习气球群区域覆盖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。