arXiv:2412.13224cs.ROcs.AI2024-12

用物理模型找最危险场景,让强化学习更安全。

Physics-model-guided Worst-case Sampling for Safe Reinforcement Learning

  • 基于物理模型主动搜索最危险的初始状态进行训练。
  • 在四足机器人上实现采样效率提升,安全策略更鲁棒。
  • 适合需要高可靠性的智能控制系统开发人员。

实际应用中学习型控制系统的事故常发生在复杂边缘情况。传统深度强化学习训练通常采用固定或均匀采样的初始状态,容易忽略关键安全场景。本文提出一种物理模型引导的最坏情况采样策略,用于训练能应对安全关键情况的强化学习策略,确保系统安全性。进一步将该策略集成到物理调控深度强化学习(Phy-DRL)框架中,构建出更高效、更安全的学习算法。通过在模拟倒立摆、2D四旋翼、模拟与真实四足机器人上的大量实验验证,新方法显著提升了采样效率,使安全策略更具鲁棒性。

原文摘要 · Abstract (English)

Real-world accidents in learning-enabled CPS frequently occur in challenging corner cases. During the training of deep reinforcement learning (DRL) policy, the standard setup for training conditions is either fixed at a single initial condition or uniformly sampled from the admissible state space. This setup often overlooks the challenging but safety-critical corner cases. To bridge this gap, this paper proposes a physics-model-guided worst-case sampling strategy for training safe policies that can handle safety-critical cases toward guaranteed safety. Furthermore, we integrate the proposed worst-case sampling strategy into the physics-regulated deep reinforcement learning (Phy-DRL) framework to build a more data-efficient and safe learning algorithm for safety-critical CPS. We validate the proposed training strategy with Phy-DRL through extensive experiments on a simulated cart-pole system, a 2D quadrotor, a simulated and a real quadruped robot, showing remarkably improved sampling efficiency to learn more robust safe policies.

强化学习安全控制物理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。