arXiv:2410.09486cs.LGcs.RO2024-10ICLR被引 27

让强化学习在保证安全的前提下高效探索,突破真实场景应用瓶颈。

ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning

  • 基于概率模型乐观规划未知动态,悲观处理安全约束。
  • 理论证明可有限时间内保证安全并逼近最优策略。
  • 适配视觉控制等高维场景,适合真实世界部署的RL研究者。

强化学习(RL)在现代AI系统中广泛应用,但现有先进方法需大量且潜在危险的环境交互才能有效学习,限制其在真实世界的应用。本文提出ActSafe,一种基于模型的新型安全强化学习算法。该算法学习系统准确的概率模型,针对未知动力学的认知不确定性进行乐观规划,同时对安全约束采取悲观策略。在动力学与约束满足一定正则性假设下,证明了算法在学习过程中可保证安全,并在有限时间内获得近似最优策略。此外,我们提出一个实用变体,结合最新的基于模型的RL进展,可在高维视觉控制任务中实现安全探索。实验证明,ActSafe在标准安全深度强化学习基准上取得当前最优性能,且学习过程全程安全。

原文摘要 · Abstract (English)

Reinforcement learning (RL) is ubiquitous in the development of modern AI systems. However, state-of-the-art RL agents require extensive, and potentially unsafe, interactions with their environments to learn effectively. These limitations confine RL agents to simulated environments, hindering their ability to learn directly in real-world settings. In this work, we present ActSafe, a novel model-based RL algorithm for safe and efficient exploration. ActSafe learns a well-calibrated probabilistic model of the system and plans optimistically w.r.t. the epistemic uncertainty about the unknown dynamics, while enforcing pessimism w.r.t. the safety constraints. Under regularity assumptions on the constraints and dynamics, we show that ActSafe guarantees safety during learning while also obtaining a near-optimal policy in finite time. In addition, we propose a practical variant of ActSafe that builds on latest model-based RL advancements and enables safe exploration even in high-dimensional settings such as visual control. We empirically show that ActSafe obtains state-of-the-art performance in difficult exploration tasks on standard safe deep RL benchmarks while ensuring safety during learning.

强化学习安全探索模型预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。