arXiv:2501.04481cs.LGcs.RO2025-01被引 1

用无监督方法生成安全数据,让智能体在少样本下也能安全探索。

Safe Reinforcement Learning with Minimal Supervision

  • 提出无监督强化学习生成初始安全数据,无需人工设计或演示。
  • 证明充足数据对学习最优安全策略至关重要,少样本时效果下降明显。
  • 适合数据稀缺但需安全探索的复杂任务,如机器人控制、自动驾驶。

现实世界中的强化学习需要确保智能体探索时不造成自身或他人伤害。现有安全强化学习方法依赖离线数据学习安全集,以实现安全在线探索,但受限于可用示范数据的数量与质量。本文研究了用于离线训练初始安全学习的數據数量与质量对在线安全强化学习能力的影响,重点关注具有空间扩展目标状态且缺乏或无示范的任务。传统方法依赖手工设计控制器或用户生成示范,成本高且难以扩展。为此,我们提出一种基于强化学习的无监督离线数据收集方法,可在无需人工控制器或示范的情况下学习复杂可扩展策略。研究结果表明,充分的示范数据对学习最优安全策略至关重要,因此我们提出乐观遗忘(optimistic forgetting)这一实用的在线安全强化学习方法,适用于数据有限场景。此外,该方法揭示了安全在线探索中多样性与最优性之间的平衡需求。

原文摘要 · Abstract (English)

Reinforcement learning (RL) in the real world necessitates the development of procedures that enable agents to explore without causing harm to themselves or others. The most successful solutions to the problem of safe RL leverage offline data to learn a safe-set, enabling safe online exploration. However, this approach to safe-learning is often constrained by the demonstrations that are available for learning. In this paper we investigate the influence of the quantity and quality of data used to train the initial safe learning problem offline on the ability to learn safe-RL policies online. Specifically, we focus on tasks with spatially extended goal states where we have few or no demonstrations available. Classically this problem is addressed either by using hand-designed controllers to generate data or by collecting user-generated demonstrations. However, these methods are often expensive and do not scale to more complex tasks and environments. To address this limitation we propose an unsupervised RL-based offline data collection procedure, to learn complex and scalable policies without the need for hand-designed controllers or user demonstrations. Our research demonstrates the significance of providing sufficient demonstrations for agents to learn optimal safe-RL policies online, and as a result, we propose optimistic forgetting, a novel online safe-RL approach that is practical for scenarios with limited data. Further, our unsupervised data collection approach highlights the need to balance diversity and optimality for safe online exploration.

安全RL无监督学习少样本数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。