arXiv:2505.00165cs.RO2025-05被引 5

用深度强化学习实现卫星姿态控制,应对推进器故障仍能精准转向。

Deep Reinforcement Learning Policies for Underactuated Satellite Attitude Control

  • 用改进的PPO算法训练神经网络策略,直接从环境交互中学习控制
  • 在正常与单个飞轮失效情况下均实现快速大角度转向和高精度定姿
  • 仿真训练的控制器可直接部署到真实硬件,具备实用潜力

自主性是未来太空探索的关键挑战。深度强化学习有望让智能体通过与环境交互自主学习复杂行为。本文研究将强化学习应用于卫星姿态控制问题,即航天器相对于惯性参考系的角向重定向。所提方法将一组控制策略实现为神经网络,使用定制版近端策略优化算法训练,使小型卫星从随机起始角度快速转向目标指向。重点考察两种工况:正常情况(三台反作用轮均工作)与欠驱动情况(随机模拟一个轴上的执行器故障)。结果表明,智能体能够有效完成大角度机动,收敛速度快,且达到工业标准的指向精度。此外,在代表性硬件上验证了该方法,证明经仿真训练的控制器在实际系统中表现良好。

原文摘要 · Abstract (English)

Autonomy is a key challenge for future space exploration endeavours. Deep Reinforcement Learning holds the promises for developing agents able to learn complex behaviours simply by interacting with their environment. This paper investigates the use of Reinforcement Learning for the satellite attitude control problem, namely the angular reorientation of a spacecraft with respect to an in- ertial frame of reference. In the proposed approach, a set of control policies are implemented as neural networks trained with a custom version of the Proximal Policy Optimization algorithm to maneuver a small satellite from a random starting angle to a given pointing target. In particular, we address the problem for two working conditions: the nominal case, in which all the actuators (a set of 3 reac- tion wheels) are working properly, and the underactuated case, where an actuator failure is simulated randomly along with one of the axes. We show that the agents learn to effectively perform large-angle slew maneuvers with fast convergence and industry-standard pointing accuracy. Furthermore, we test the proposed method on representative hardware, showing that by taking adequate measures controllers trained in simulation can perform well in real systems.

强化学习卫星控制欠驱动系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。