用强化学习提前决策卫星避撞,省燃料还更安全。
A Markov Decision Process Framework for Early Maneuver Decisions in Satellite Collision Avoidance
- 构建连续状态离散动作的马尔可夫决策模型,提前决定避撞时机。
- 相比传统24小时前决策,合成数据上总耗燃料减少显著,单次操作更省。
- 适合关注卫星寿命与轨道安全的航天任务规划人员。
我们提出一种基于马尔可夫决策过程(MDP)的自主卫星避撞机动(CAM)引导决策框架,并采用强化学习策略梯度(RL-PG)算法,直接利用历史CAM数据优化引导策略。该方法在保证可接受碰撞风险的前提下,通过提前决策以最小化平均推进剂消耗。将CAM建模为连续状态、离散动作、有限时域的MDP,关键决策在于何时启动机动。MDP通过解析模型综合碰撞风险、推进剂消耗和转移轨道几何形状来定义奖励。相比传统在距离最近接近点(TCA)前24小时触发的截断策略,所训练策略在合成关联事件中整体及单次机动耗能均显著降低;在历史关联事件中虽总耗能略高,但单次耗能更低。该策略对需采取机动的事件判定稍显保守,但整体提升了安全性与效率。
原文摘要 · Abstract (English)
We develop a Markov decision process (MDP) framework to autonomously make guidance decisions for satellite collision avoidance maneuver (CAM) and a reinforcement learning policy gradient (RL-PG) algorithm to enable direct optimization of guidance policy using historic CAM data. In addition to maintaining acceptable collision risks, this approach seeks to minimize the average propellant consumption of CAMs by making early maneuver decisions. We model CAM as a continuous state, discrete action and finite horizon MDP, where the critical decision is determining when to initiate the maneuver. The MDP models decision rewards using analytical models of collision risk, propellant consumption, and transit orbit geometry. By deciding to maneuver earlier than conventional methods, the Markov policy effectively favors CAMs that achieve comparable rates of collision risk reduction while consuming less propellant. Using historical data of tracked conjunction events, we verify this framework and conduct an extensive parameter-sensitivity study. When evaluated on synthetic conjunction events, the trained policy consumes significantly less propellant overall and per maneuver in comparison to a conventional cut-off policy that initiates maneuvers 24 hours before the time of closest approach (TCA). On historical conjunction events, the trained policy consumes more propellant overall but consumes less propellant per maneuver. For both historical and synthetic conjunction events, the trained policy is slightly more conservative in identifying conjunctions events that warrant CAMs in comparison to cutoff policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。