用深度强化学习提升航天器再入姿态控制精度与鲁棒性
Deep Reinforcement Learning for Spacecraft Attitude Control During Atmospheric Re-Entry
- 采用连续离线策略强化学习,结合动态随机化提升泛化能力
- 在质量、惯量、舵机带宽变化下表现优于传统控制器
- 适合对复杂环境适应性要求高的航天器自主控制系统
深度强化学习有望通过更有效地处理非线性动力学、不确定性及故障情况,实现比传统方法更自适应、精确和鲁棒的姿态控制。本文探索了强化学习在航天器再入过程中的姿态控制应用。以工业标准的变增益比例-积分-微分控制器作为基线,评估了无模型强化学习及混合控制器的表现。采用连续、离线策略强化学习框架,最先进的RL方法在该任务中达到与传统控制相当的性能。然而其分布外泛化能力不足,因此训练中引入动态随机化以模拟挑战性任务变化,并在预定义操作包络内强制泛化。最终评估表明,最优强化学习控制器在应用特定指标下表现更优:混合控制器能更准确跟踪攻角,且在质量、惯量张量及舵机带宽变化下更具鲁棒性。
原文摘要 · Abstract (English)
Deep reinforcement learning has the potential to solve attitude control problems more adaptively, precisely, and robustly by handling nonlinear dynamics, uncertainties, and failure cases more effectively than traditional attitude control approaches. We explore reinforcement learning (RL) for attitude control in spacecraft re-entry. An industry-standard proportional-integral-derivative controller with gain scheduling serves as a strong baseline for model-free RL and hybrid controllers that combine these two approaches. We formalize the application in the RL framework to apply continuous, off-policy RL. State-of-the-art RL achieves comparable performance to traditional control approaches in this domain. However, its out-of-distribution generalization is not sufficient. Hence, we use dynamics randomization to introduce challenging task variations during training and enforce generalization in a predefined operational envelope. Finally, we assess the best obtained RL-based controllers with application-specific metrics to show superior performance in comparison to traditional controllers in the operational envelope, that is, hybrid controllers are able to track the angle of attack better and are more robust under variations of mass, inertia tensor, and flap actuator bandwidth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。