arXiv:2504.04160cs.LGcs.MA2025-04NeurIPS被引 2

构建高保真卫星多智能体环境,提升强化学习在轨自主决策能力

OrbitZoo: Real Orbital Systems Challenges for Reinforcement Learning

  • 基于工业级动力学库构建真实轨道环境,支持多智能体协同
  • 与星链星座对比误差仅0.16%,验证了动力学精度可靠性
  • 适合研究太空碰撞规避与自主轨道控制的科研人员

卫星与轨道碎片数量激增导致太空拥堵问题日益严重,威胁卫星安全与可持续性。碰撞规避、轨道保持及机动等任务需应对动态不确定性与多智能体交互。强化学习(RL)在该领域展现出潜力,可实现自适应、自主的太空操作策略;但现有多数框架依赖从零构建的简化仿真环境,需大量时间实现和验证轨道动力学,难以充分反映真实复杂性。为此,我们提出OrbitZoo,一个基于高保真工业标准库的多智能体强化学习环境,支持真实数据生成与多种场景如碰撞规避与协同机动,确保精确可靠的轨道动力学建模。该环境经真实星链星座验证,与实际数据对比的平均绝对百分比误差(MAPE)仅为0.16%。此精度保障了高保真仿真的可靠性,助力实现独立自主的卫星运行。

原文摘要 · Abstract (English)

The increasing number of satellites and orbital debris has made space congestion a critical issue, threatening satellite safety and sustainability. Challenges such as collision avoidance, station-keeping, and orbital maneuvering require advanced techniques to handle dynamic uncertainties and multi-agent interactions. Reinforcement learning (RL) has shown promise in this domain, enabling adaptive, autonomous policies for space operations; however, many existing RL frameworks rely on custom-built environments developed from scratch, which often use simplified models and require significant time to implement and validate the orbital dynamics, limiting their ability to fully capture real-world complexities. To address this, we introduce OrbitZoo, a versatile multi-agent RL environment built on a high-fidelity industry standard library, that enables realistic data generation, supports scenarios like collision avoidance and cooperative maneuvers, and ensures robust and accurate orbital dynamics. The environment is validated against a real satellite constellation, Starlink, achieving a Mean Absolute Percentage Error (MAPE) of 0.16% compared to real-world data. This validation ensures reliability for generating high-fidelity simulations and enabling autonomous and independent satellite operations.

强化学习轨道控制多智能体仿真环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。