arXiv:2509.21045cs.ROcs.LG2025-09被引 4

融合强化学习与模型预测控制,实现带燃料晃动的卫星自主对接

MPC-based Deep Reinforcement Learning Method for Space Robotic Control with Fuel Sloshing Mitigation

  • 用MPC预测动态,加速强化学习训练并提升控制鲁棒性
  • 在6自由度对接中实现更高精度、成功率和更低控制能耗
  • 适合航天器自主对接与在轨加注任务研究者参考

本文提出一种融合强化学习(RL)与模型预测控制(MPC)的综合框架,用于微重力环境下部分充液燃料罐卫星的自主对接。传统对接受燃料晃动影响,产生不可预测力矩,危及稳定性。为此,将近端策略优化(PPO)与软动作价值批判(SAC)算法与MPC结合,利用MPC的预测能力加速RL训练并增强控制鲁棒性。通过零重力实验室(Zero-G Lab of SnT)平面稳定实验及高保真数值仿真,验证了6-DOF对接中燃料晃动动力学下的性能。仿真结果表明,SAC-MPC方法在对接精度、成功率达98%以上、控制能耗降低30%以上方面均优于独立RL及PPO-MPC方法。本研究推动了燃料高效、抗扰动的卫星对接技术发展,提升了在轨加注与服务任务的可行性。

原文摘要 · Abstract (English)

This paper presents an integrated Reinforcement Learning (RL) and Model Predictive Control (MPC) framework for autonomous satellite docking with a partially filled fuel tank. Traditional docking control faces challenges due to fuel sloshing in microgravity, which induces unpredictable forces affecting stability. To address this, we integrate Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) RL algorithms with MPC, leveraging MPC's predictive capabilities to accelerate RL training and improve control robustness. The proposed approach is validated through Zero-G Lab of SnT experiments for planar stabilization and high-fidelity numerical simulations for 6-DOF docking with fuel sloshing dynamics. Simulation results demonstrate that SAC-MPC achieves superior docking accuracy, higher success rates, and lower control effort, outperforming standalone RL and PPO-MPC methods. This study advances fuel-efficient and disturbance-resilient satellite docking, enhancing the feasibility of on-orbit refueling and servicing missions.

强化学习航天控制燃料晃动模型预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。