用闭式解解决多约束安全强化学习的计算难题
Multi-Constraint Safe Reinforcement Learning via Closed-form Solution for Log-Sum-Exp Approximation of Control Barrier Functions
- 用单个复合CBF近似多个安全约束的连续且可微逻辑
- 推导出策略网络的闭式解,避免端到端优化开销
- 适合需要理论安全保障的机器人控制场景
强化学习(RL)在训练任务策略及其部署中的安全性已成为安全强化学习领域的核心关注点。尽管基于控制屏障函数(CBF)的安全策略在一系列控制仿射型机器人系统中表现成功,但将其与强化学习结合仍面临挑战。首先,将安全优化嵌入强化学习训练流程需保证优化输出对输入参数可微,即要求可微优化,而该问题难以求解;其次,现有可微优化框架在处理多约束问题时存在显著效率瓶颈。为此,本文提出一种基于CBF的安全强化学习架构,有效应对上述问题。该方法通过单一复合CBF构建多约束的连续且可微的逻辑与(AND)近似,进而推导出适用于强化学习策略网络的二次规划闭式解,从而避免在端到端安全强化学习管道中进行可微优化。该策略因采用闭式解,大幅降低计算复杂度,同时保持理论安全保证。仿真结果表明,相比依赖可微优化的现有方法,所提方法显著降低训练计算成本,并在整个训练过程中确保可证明的安全性。
原文摘要 · Abstract (English)
The safety of training task policies and their subsequent application using reinforcement learning (RL) methods has become a focal point in the field of safe RL. A central challenge in this area remains the establishment of theoretical guarantees for safety during both the learning and deployment processes. Given the successful implementation of Control Barrier Function (CBF)-based safety strategies in a range of control-affine robotic systems, CBF-based safe RL demonstrates significant promise for practical applications in real-world scenarios. However, integrating these two approaches presents several challenges. First, embedding safety optimization within the RL training pipeline requires that the optimization outputs be differentiable with respect to the input parameters, a condition commonly referred to as differentiable optimization, which is non-trivial to solve. Second, the differentiable optimization framework confronts significant efficiency issues, especially when dealing with multi-constraint problems. To address these challenges, this paper presents a CBF-based safe RL architecture that effectively mitigates the issues outlined above. The proposed approach constructs a continuous AND logic approximation for the multiple constraints using a single composite CBF. By leveraging this approximation, a close-form solution of the quadratic programming is derived for the policy network in RL, thereby circumventing the need for differentiable optimization within the end-to-end safe RL pipeline. This strategy significantly reduces computational complexity because of the closed-form solution while maintaining safety guarantees. Simulation results demonstrate that, in comparison to existing approaches relying on differentiable optimization, the proposed method significantly reduces training computational costs while ensuring provable safety throughout the training process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。