提出一种安全强化学习方法,实现持续安全探索。
Sampling-Based Safe Reinforcement Learning

- 基于动态采样构建约束,确保学习全过程安全
- 在连续空间中实现高概率安全保证和近优策略的快速收敛
- 无需额外探索奖励,适合真实机器人部署
安全探索仍是强化学习中的根本挑战,限制了其在现实世界中的应用。我们提出基于采样的安全强化学习(SBSRL),一种模型基强化学习算法,通过在有限动态样本集上联合施加约束,维持学习过程中的安全性。该方法近似处理不确定动态下的不可行最坏情况优化,在连续域中实现实用的安全保障。我们进一步提出一种基于认知不确定性约束的探索策略,无需显式探索奖励。在常规条件下,我们推导出学习全程的高概率安全保证及恢复近优策略的有限时间样本复杂度界。实验表明,SBSRL在仿真与真实机器人硬件上均实现安全高效的探索,并可扩展至高维连续控制问题的深度集成实现。
原文摘要 · Abstract (English)
Safe exploration remains a fundamental challenge in reinforcement learning (RL), limiting the deployment of RL agents in the real world. We propose Sampling-Based Safe Reinforcement Learning (SBSRL), a model-based RL algorithm that maintains safety throughout the learning process by enforcing constraints jointly across a finite set of dynamics samples. This formulation approximates an intractable worst-case optimization over uncertain dynamics and enables practical safety guarantees in continuous domains. We further introduce an exploration strategy based on constraining epistemic uncertainty, eliminating the need for explicit exploration bonuses. Under regularity conditions, we derive high-probability guarantees of safety throughout learning and a finite-time sample complexity bound for recovering a near-optimal policy. Empirically, SBSRL achieves safe and efficient exploration both in simulation and in real robotic hardware, and readily extends to practical deep-ensemble implementations that scale to high-dimensional continuous control problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。