用强化学习让模块化卫星自动规划多种构型路径,成功率超95%。
A Goal-Oriented Reinforcement Learning-Based Path Planning Algorithm for Modular Self-Reconfigurable Satellites
- 基于目标导向的强化学习,支持多目标构型路径规划
- 四单元和六单元集群路径成功率分别达95%和73%
- 融合事后经验回放与无效动作掩码,解决稀疏奖励难题
模块化可重构卫星由多个可自主调整构型的模块单元组成,构型变化使其能执行多样任务。现有重构路径规划算法普遍存在计算复杂度高、泛化能力差、难以支持多种目标构型的问题。本文提出一种面向目标的强化学习路径规划算法,首次解决以往强化学习方法无法处理多目标构型的挑战。通过引入事后经验回放(Hindsight Experience Replay)和无效动作掩码(Invalid Action Masking)技术,有效克服稀疏奖励与无效动作带来的困难。实验表明,该模型在由4个和6个模块组成的卫星集群中,对任意目标构型的路径规划成功率分别达到95%和73%。
原文摘要 · Abstract (English)
Modular self-reconfigurable satellites refer to satellite clusters composed of individual modular units capable of altering their configurations. The configuration changes enable the execution of diverse tasks and mission objectives. Existing path planning algorithms for reconfiguration often suffer from high computational complexity, poor generalization capability, and limited support for diverse target configurations. To address these challenges, this paper proposes a goal-oriented reinforcement learning-based path planning algorithm. This algorithm is the first to address the challenge that previous reinforcement learning methods failed to overcome, namely handling multiple target configurations. Moreover, techniques such as Hindsight Experience Replay and Invalid Action Masking are incorporated to overcome the significant obstacles posed by sparse rewards and invalid actions. Based on these designs, our model achieves a 95% and 73% success rate in reaching arbitrary target configurations in a modular satellite cluster composed of four and six units, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。