arXiv:2410.11671cs.ROcs.LG2024-10被引 26

让强化学习模型在训练时就学会配合安全过滤器,提升性能与效率。

Safety Filtering While Training: Improving the Performance and Sample Efficiency of Reinforcement Learning Agents

  • 在训练阶段融合安全过滤器,使智能体提前适应其约束。
  • 实测显示性能提升,样本效率提高,且减少控制抖动现象。
  • 适合关注安全强化学习的科研人员与工程实践者。

强化学习控制器灵活高效,但通常无法保证安全性。安全过滤器能为强化学习控制器提供硬性安全保障,同时保持灵活性。然而,由于控制器与安全过滤器分离,可能导致不期望的行为,常导致性能下降和鲁棒性降低。本文分析了将安全过滤器融入强化学习训练过程的多种改进方法,而非仅在评估阶段应用。这些修改使强化学习控制器学会适应安全过滤器,从而提升整体表现。论文通过仿真和真实世界实验(使用Crazyflie 2.0无人机)进行了全面评估,考察不同训练策略与超参数对性能、样本效率、安全性及抖动(chattering)的影响。研究结果可为安全强化学习领域的研究人员与实践者提供重要指导。

原文摘要 · Abstract (English)

Reinforcement learning (RL) controllers are flexible and performant but rarely guarantee safety. Safety filters impart hard safety guarantees to RL controllers while maintaining flexibility. However, safety filters can cause undesired behaviours due to the separation between the controller and the safety filter, often degrading performance and robustness. In this paper, we analyze several modifications to incorporating the safety filter in training RL controllers rather than solely applying it during evaluation. The modifications allow the RL controller to learn to account for the safety filter, improving performance. This paper presents a comprehensive analysis of training RL with safety filters, featuring simulated and real-world experiments with a Crazyflie 2.0 drone. We examine how various training modifications and hyperparameters impact performance, sample efficiency, safety, and chattering. Our findings serve as a guide for practitioners and researchers focused on safety filters and safe RL.

强化学习安全控制无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。