arXiv:2505.18858cs.ROcs.LG2025-05被引 1

用安全约束函数指导机器人学习,让强化学习更安全可靠。

Guided by Guardrails: Control Barrier Functions as Safety Instructors for Robotic Learning

  • 引入连续负奖励模拟长期安全风险,替代传统即时惩罚。
  • 结合控制屏障函数,使机器人在仿真和真实场景中避免危险区域。
  • 适合希望提升机器人学习安全性的研究者与工程师。

安全是阻碍基于学习的机器人系统在日常生活中广泛应用的主要障碍。尽管强化学习(RL)作为机器人学习范式展现出潜力,但传统RL框架通常通过单个标量负奖励并立即终止任务来建模安全问题,无法捕捉不安全行为的时序后果(如持续碰撞损伤)。本文提出一种新方法,通过施加不终止任务的连续负奖励来模拟这些时序效应。实验表明,标准RL方法在此模型下表现不佳,因为不安全区域累积的负值形成学习障碍。为此,我们证明了控制屏障函数(CBFs)——因其已知的安全保障——能有效引导机器人避开灾难性区域,并提升学习效果。本文提出三种基于CBF的融合方法,将传统RL与CBF结合,指导智能体学习安全行为。在模拟环境及使用四轮差速驱动机器人的真实场景中进行的实证分析,验证了这些方法在实现安全机器人学习方面的可行性。

原文摘要 · Abstract (English)

Safety stands as the primary obstacle preventing the widespread adoption of learning-based robotic systems in our daily lives. While reinforcement learning (RL) shows promise as an effective robot learning paradigm, conventional RL frameworks often model safety by using single scalar negative rewards with immediate episode termination, failing to capture the temporal consequences of unsafe actions (e.g., sustained collision damage). In this work, we introduce a novel approach that simulates these temporal effects by applying continuous negative rewards without episode termination. Our experiments reveal that standard RL methods struggle with this model, as the accumulated negative values in unsafe zones create learning barriers. To address this challenge, we demonstrate how Control Barrier Functions (CBFs), with their proven safety guarantees, effectively help robots avoid catastrophic regions while enhancing learning outcomes. We present three CBF-based approaches, each integrating traditional RL methods with Control Barrier Functions, guiding the agent to learn safe behavior. Our empirical analysis, conducted in both simulated environments and real-world settings using a four-wheel differential drive robot, explores the possibilities of employing these approaches for safe robotic learning.

强化学习安全控制机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。