arXiv:2602.17685cs.LGcs.RO2026-02被引 1

用深度强化学习优化低轨多目标太空垃圾清理,兼顾效率与安全。

Optimal Multi-Debris Mission Planning in LEO: A Deep Reinforcement Learning Approach with Co-Elliptic Transfers and Refueling

  • 采用共椭圆转移结合油料补给的统一机动框架
  • 强化学习比贪心法多清理两倍垃圾,运行更快
  • 适合研究空间自主任务规划与在轨维护的学者

本文针对低地球轨道(LEO)多目标主动空间垃圾清除(ADR)问题,提出一种统一的共椭圆机动框架,整合霍曼转移、安全椭圆接近操作及显式油料补给逻辑。在包含随机分布垃圾场、禁飞区和ΔV约束的真实轨道仿真环境中,对比了三种规划算法:贪心启发式、蒙特卡洛树搜索(MCTS)与基于掩码近端策略优化(Masked PPO)的深度强化学习。100次测试结果显示,Masked PPO在任务效率与计算性能上均表现最优,清理的垃圾数量是贪心法的两倍,且显著优于MCTS的运行时间。结果表明,现代强化学习方法在可扩展、安全且资源高效的太空任务规划中具有巨大潜力,为未来主动垃圾清除自主化提供支持。

原文摘要 · Abstract (English)

This paper addresses the challenge of multi target active debris removal (ADR) in Low Earth Orbit (LEO) by introducing a unified coelliptic maneuver framework that combines Hohmann transfers, safety ellipse proximity operations, and explicit refueling logic. We benchmark three distinct planning algorithms Greedy heuristic, Monte Carlo Tree Search (MCTS), and deep reinforcement learning (RL) using Masked Proximal Policy Optimization (PPO) within a realistic orbital simulation environment featuring randomized debris fields, keep out zones, and delta V constraints. Experimental results over 100 test scenarios demonstrate that Masked PPO achieves superior mission efficiency and computational performance, visiting up to twice as many debris as Greedy and significantly outperforming MCTS in runtime. These findings underscore the promise of modern RL methods for scalable, safe, and resource efficient space mission planning, paving the way for future advancements in ADR autonomy.

空间任务规划强化学习垃圾清理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。