用强化学习优化太空碎片巡检顺序,显著缩短任务时间。
Revisiting Space Mission Planning: A Reinforcement Learning-Guided Approach for Multi-Debris Rendezvous
- 采用掩码PPO算法学习最优碎片访问序列。
- 相比遗传和贪心算法,任务总时长平均减少10.96%和13.66%。
- 适合需要高效规划的太空碎片清除任务研究者。
本研究将深度强化学习中的掩码近端策略优化(PPO)算法引入太空碎片任务规划,通过伊佐(Izzo)改进的Lambert求解器计算每次交会机动,旨在确定所有给定碎片的最优访问顺序,以最小化整项任务的总时间。开发了神经网络策略,在不同碎片分布的模拟任务中进行训练。训练后,该模型利用伊佐版本的Lambert机动快速计算近似最优路径。与标准启发式方法对比,强化学习方法在规划效率上表现突出:相比遗传算法和贪心算法,平均任务时间分别减少约10.96%和13.66%。模型在多种模拟场景下均能快速识别最省时的碎片访问序列,展现出卓越的计算速度与适应性,标志着空间碎片清除任务规划策略的重要进展。
原文摘要 · Abstract (English)
This research introduces a novel application of a masked Proximal Policy Optimization (PPO) algorithm from the field of deep reinforcement learning (RL), for determining the most efficient sequence of space debris visitation, utilizing the Lambert solver as per Izzo's adaptation for individual rendezvous. The aim is to optimize the sequence in which all the given debris should be visited to get the least total time for rendezvous for the entire mission. A neural network (NN) policy is developed, trained on simulated space missions with varying debris fields. After training, the neural network calculates approximately optimal paths using Izzo's adaptation of Lambert maneuvers. Performance is evaluated against standard heuristics in mission planning. The reinforcement learning approach demonstrates a significant improvement in planning efficiency by optimizing the sequence for debris rendezvous, reducing the total mission time by an average of approximately {10.96\%} and {13.66\%} compared to the Genetic and Greedy algorithms, respectively. The model on average identifies the most time-efficient sequence for debris visitation across various simulated scenarios with the fastest computational speed. This approach signifies a step forward in enhancing mission planning strategies for space debris clearance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。