用元强化学习优化航天器近距离操作的安全控制,省燃料还更安全
Memory-Efficient Meta-Reinforcement Learning for Adaptive Safety-Critical Control in Adversarial Spacecraft Proximity Operations

- 用LSTM、GRU和Mamba等循环网络,结合PPO/SAC算法训练安全函数
- Mamba+PPO在合作与对抗场景下完成率超95%,燃料节省超30%
- 适合做高可靠性航天器自主控制的工程师和研究人员
自主航天器交会与近距离操作(RPO)需要在推力约束下保证安全并最小化燃料消耗。输入约束控制屏障函数(ICCBF)为带有执行器约束的非线性系统提供了一种构造前向不变安全集的控制方法。先前研究显示,通过元强化学习(meta-RL)学习定义ICCBF递归的类-𝐾函数,可实现鲁棒且非贪婪的安全关键控制。本文进一步研究了三种循环网络架构(LSTM、GRU、选择性状态空间模型Mamba)和两种训练算法(PPO、SAC),以确定最优的ICCBF类-𝐾函数元-RL调优方案。除合作测试外,还在对抗行为场景下评估性能,即目标航天器故意恶化追踪航天器的安全状况。结果表明,在所有测试的协作与非协作场景中,采用PPO的Mamba状态空间模型在任务完成率、安全性及燃料节约方面均优于其他架构,表现显著领先。
原文摘要 · Abstract (English)
Autonomous spacecraft rendezvous and proximity operations (RPO) require controllers that guarantee safety under thrust constraints while minimizing fuel expenditure. Input-constrained control barrier functions (ICCBFs) provide a control method for nonlinear systems with actuation constraints that construct a forward-invariant safe set. Previous work has shown that learning class-$\mathcal{K}$ functions defining the ICCBF recursion via meta reinforcement learning (meta-RL) yields a robust, non-greedy approach to safety-critical control in RPO. This paper extends that framework further by investigating the performance of three recurrent network architectures (Long Short Term Memory (LSTM), Gated Recurrent Unit (GRU), Selective State Space Model (Mamba)) and two training algorithms (Proximal Policy Optimization (PPO) and Soft Actor Critic (SAC)) to identify the best setup for tuning ICCBF class-K functions via meta-RL. In addition to cooperative test cases, performance is evaluated in the presence of adversarial behavior where the target spacecraft behaves in a way that worsens the safety of the chaser spacecraft. Results indicate that state space models such as Mamba when used with PPO achieve superior task completion, safety, and fuel-savings compared to other architectures, across all cooperative and uncooperative scenarios tested.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。