arXiv:2608.09628cs.LGcs.RO2026-08被引 2

用强化学习自动避让太空碎片,成功率超97%

Satellite Trajectory Optimization via Proximal Policy Optimization for Space Debris Avoidance

论文配图:Satellite Trajectory Optimization via Proximal Policy Optimization for Space Debris Avoidance
图 1 · 摘自论文原文
  • 基于近端策略优化训练智能体,实现自主避障
  • 1000次测试中避障成功率达97.5%,远超传统方法
  • 开源仿真平台支持真实碎片环境训练,适合航天安全研究者

低地球轨道(LEO)和地球同步轨道(GEO)的碰撞规避系统正面临日益增长的碎片事件威胁,尤其受巨型星座发射加剧轨道拥堵影响。当前多为人工或规则驱动的避障方式,难以适应动态变化环境。为此,本文提出一种基于近端策略优化(PPO)的强化学习避障策略,并在开源高保真天体动力学仿真器上训练与评估。在1000次确定性GEO场景测试中,该智能体实现97.5%的避障成功率,显著优于规则基基准(20.7%)和脉冲速度增量规划基准(27.5%)。仿真器融合牛顿二体动力学、太阳/月球三体扰动、燃料依赖推力及可配置碎片场。训练采用课程学习与奖励塑形,目标为提升生存率、足够最小接近距离及Δv节约。评估采用完全确定性流程,包含共享随机种子、每集日志及遥测导出。相关框架已公开:https://purl.org/sat-trajectory-avoidance

原文摘要 · Abstract (English)

Collision avoidance systems are commonly used to avoid fragmentation events occurring in Low-Earth Orbit (LEO) and Geosynchronous Equatorial Orbit (GEO). However, these events have been growing in frequency as orbital congestion worsens with the launch of megaconstellations. Consequently, conjunction alerts and collision risks are becoming increasingly common. Current practices, which are commonly manual or rule-based, have difficulty scaling to these worsening dynamic environments. To address this intensifying situation, we propose a reinforcement-learning policy for autonomous collision avoidance, trained via Proximal Policy Optimization (PPO) along with an open-source, high-fidelity astrodynamics simulator for training and evaluation. In 1,000 deterministic GEO episodes, our agent achieves a 97.5% collision avoidance success rate, outperforming traditional controllers such as a rule-based baseline (20.7% success) and an impulsive delta-v planner baseline (27.5% success). To achieve these results, we designed a simulator to train and evaluate our agent, using real-world and simulated debris. We simulate Newtonian two-body dynamics using Sun/Moon third-body perturbations, fuel-dependent thrust, and configurable debris fields. The agent is trained with curriculum learning and shaped rewards oriented toward encouraging survival, adequate projected miss distance, and delta-v conservation. Finally, our evaluation consisted of a fully deterministic pipeline, including shared seeds, per-episode logs, and telemetry exports. Our work is a publicly available framework at https://purl.org/sat-trajectory-avoidance

强化学习太空避障PPO轨道优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。