利用离线数据学习技能风险,实现更安全高效的强化学习。
Skill-based Safe Reinforcement Learning with Risk Planning
- 用离线演示数据训练技能风险预测器
- 在线环境中通过风险规划提升安全策略性能
- 适合需高安全性的机器人控制场景
安全强化学习(Safe RL)旨在确保智能体在与真实世界环境交互时保持安全,避免因不当动作导致高成本或严重后果。本文提出一种新型安全技能规划方法(SSkP),通过利用离线演示数据来增强有效安全RL。该方法分为两阶段:首先使用PU学习从离线演示数据中学习技能风险预测器;随后基于该预测器设计新的风险规划机制,在在线环境中高效学习风险规避的安全策略,同时动态适应环境并更新风险预测器。我们在多个基准机器人仿真环境中进行实验,结果表明所提方法持续优于现有最先进安全RL方法。
原文摘要 · Abstract (English)
Safe Reinforcement Learning (Safe RL) aims to ensure safety when an RL agent conducts learning by interacting with real-world environments where improper actions can induce high costs or lead to severe consequences. In this paper, we propose a novel Safe Skill Planning (SSkP) approach to enhance effective safe RL by exploiting auxiliary offline demonstration data. SSkP involves a two-stage process. First, we employ PU learning to learn a skill risk predictor from the offline demonstration data. Then, based on the learned skill risk predictor, we develop a novel risk planning process to enhance online safe RL and learn a risk-averse safe policy efficiently through interactions with the online RL environment, while simultaneously adapting the skill risk predictor to the environment. We conduct experiments in several benchmark robotic simulation environments. The experimental results demonstrate that the proposed approach consistently outperforms previous state-of-the-art safe RL methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。