零样本安全强化学习让复杂系统在不训练全模型下保持安全。
Safety Guarantees in Zero-Shot Reinforcement Learning for Cascade Dynamical Systems

- 用简化模型训练安全策略,忽略内部状态动态但将其视为影响外部的输入。
- 理论证明部署后系统保持安全的概率与底层控制器跟踪能力相关。
- 适合关注复杂系统安全控制的工程师和研究者参考。
本文研究级联动力系统中的零样本安全保证问题。这类系统中,部分状态(内态)影响其余状态(外态)的演化,但反之不成立。安全性定义为系统在所有时间都以高概率停留在安全集合内。我们提出在降阶模型上训练安全强化学习策略,该模型忽略内态动态,但将其视为影响外态的动作输入,从而降低训练复杂度。部署时,训练好的策略与低层控制器结合,由其跟踪策略提供的参考轨迹。主要理论贡献是给出了全阶系统中保持安全的概率上界,揭示了零样本部署后的安全概率与内态跟踪质量之间的关系。在四旋翼导航任务上的实验验证表明,安全保证能否维持取决于底层控制器的带宽和跟踪性能。
原文摘要 · Abstract (English)
This paper considers the problem of zero-shot safety guarantees for cascade dynamical systems. These are systems where a subset of the states (the inner states) affects the dynamics of the remaining states (the outer states) but not vice-versa. We define safety as remaining on a set deemed safe for all times with high probability. We propose to train a safe RL policy on a reduced-order model, which ignores the dynamics of the inner states, but it treats it as an action that influences the outer state. Thus, reducing the complexity of the training. When deployed in the full system the trained policy is combined with a low-level controller whose task is to track the reference provided by the RL policy. Our main theoretical contribution is a bound on the safe probability in the full-order system. In particular, we establish the interplay between the probability of remaining safe after the zero-shot deployment and the quality of the tracking of the inner states. We validate our theoretical findings on a quadrotor navigation task, demonstrating that the preservation of the safety guarantees is tied to the bandwidth and tracking capabilities of the low-level controller.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。