用强化学习优化小卫星多目标清轨任务,兼顾省油与避障。
Optimizing Mission Planning for Multi-Debris Rendezvous Using Reinforcement Learning with Refueling and Adaptive Collision Avoidance
- 用掩码PPO算法让卫星实时调整轨道,动态避障并规划路径。
- 相比传统方法,碰撞风险更低,任务耗时和燃料消耗均减少。
- 适合研究自主空间任务规划或太空清障的科研人员参考。
随着地球轨道上碎片日益增多,主动碎片清除(ADR)任务面临安全运行与规避碰撞的巨大挑战。本文提出一种基于强化学习(RL)的框架,用于提升小卫星执行多目标碎片清除任务时的自适应避障能力。小卫星因灵活性强、成本低、机动性好,特别适合此类动态任务。该框架整合了补给策略、高效任务规划与自适应避障机制,采用掩码近端策略优化(masked PPO)算法,使代理能根据实时轨道状况动态调整机动行为。关键考量包括燃料效率、避开高风险碰撞区域以及优化动态轨道参数。模型学习生成多碎片目标的最优交会序列,在保证安全性的同时降低燃料消耗与任务时间,并合理安排补给节点。通过基于Iridium 33碎片数据集的模拟场景验证,涵盖多种轨道构型与碎片分布,结果表明该方法在降低碰撞风险的同时显著提升任务效率。本研究为复杂多目标清轨任务提供了可扩展的解决方案,适用于其他自主空间任务规划问题。
原文摘要 · Abstract (English)
As the orbital environment around Earth becomes increasingly crowded with debris, active debris removal (ADR) missions face significant challenges in ensuring safe operations while minimizing the risk of in-orbit collisions. This study presents a reinforcement learning (RL) based framework to enhance adaptive collision avoidance in ADR missions, specifically for multi-debris removal using small satellites. Small satellites are increasingly adopted due to their flexibility, cost effectiveness, and maneuverability, making them well suited for dynamic missions such as ADR. Building on existing work in multi-debris rendezvous, the framework integrates refueling strategies, efficient mission planning, and adaptive collision avoidance to optimize spacecraft rendezvous operations. The proposed approach employs a masked Proximal Policy Optimization (PPO) algorithm, enabling the RL agent to dynamically adjust maneuvers in response to real-time orbital conditions. Key considerations include fuel efficiency, avoidance of active collision zones, and optimization of dynamic orbital parameters. The RL agent learns to determine efficient sequences for rendezvousing with multiple debris targets, optimizing fuel usage and mission time while incorporating necessary refueling stops. Simulated ADR scenarios derived from the Iridium 33 debris dataset are used for evaluation, covering diverse orbital configurations and debris distributions to demonstrate robustness and adaptability. Results show that the proposed RL framework reduces collision risk while improving mission efficiency compared to traditional heuristic approaches. This work provides a scalable solution for planning complex multi-debris ADR missions and is applicable to other multi-target rendezvous problems in autonomous space mission planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。