用模型预测控制保障强化学习安全,实现物理系统可靠探索。
Safe Reinforcement Learning using Ideas from Model Predictive Control

- 结合深度强化学习与模型预测控制,构建安全约束的动态规划框架。
- 在非线性单自由度实验平台验证,成功实现安全策略收敛与硬件部署。
- 适合需严格安全约束的机器人与复杂物理系统控制场景。
强化学习(RL)可直接从数据中合成控制策略,对复杂网络物理系统(CPS)和机器人极具吸引力。然而,一个长期挑战是确保在主动学习阶段满足严格的硬性安全约束。在真实物理系统中,违反机械极限可能导致不可逆损坏,因此探索必须严格限制在安全操作区域内。本文提出一种通用框架,将深度强化学习(DRL)的自适应、高性能特性与模型预测控制(MPC)的形式化安全保证相结合。利用系统动力学的数学模型,离线计算生成可行状态-动作空间,代表所有能保证约束满足的安全组合。训练与部署过程中,通过安全滤波器将强化学习代理的瞬时动作投影至该全局验证的可行集中。我们在一个非线性1-DoF实验室测试平台系统评估该方法,证明了其在物理硬件上的成功探索与稳定策略收敛。
原文摘要 · Abstract (English)
Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems (CPSs) and robotics. A persistent challenge, however, is ensuring strict, hard safety constraints during the active learning phase. In real-world physical systems, violating mechanical limits can cause irreversible damage, necessitating that exploration remains strictly within safe operational regions. We propose a generalized framework that combines the adaptive, high-performance nature of deep reinforcement learning (DRL) with the formal safety guarantees of model predictive control (MPC). Using a mathematical model of the system dynamics, offline MPC computations define a feasible state-action space, representing all safe combinations of system states and control inputs that guarantee constraint satisfaction. During training and deployment, the RL agent's instantaneous actions are projected onto this globally verified feasible set via a safety filter. We systematically evaluate our generalized approach on a non-linear 1-DoF laboratory testbed, demonstrating successful exploration and stable policy convergence on physical hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。