用强化学习拓展MPC路径规划,提升自动驾驶安全与性能。
Safety Reinforced Model Predictive Control (SRMPC): Improving MPC with Reinforcement Learning for Motion Planning in Autonomous Driving
- 结合安全强化学习与MPC,突破传统凸近似限制
- 在高速场景中同时提升安全性与路径性能
- 适合关注自动驾驶规划安全性的研究者
模型预测控制(MPC)广泛用于自动驾驶路径规划,其实时性依赖于对最优控制问题(OCP)的凸近似,但此类近似将解限制在局部子空间,可能无法找到全局最优。为此,我们提出安全强化学习(SRL)驱动的MPC框架,在MPC内部生成新的安全参考轨迹。通过学习策略,MPC可跳出前一解的邻域,探索更优解。采用约束强化学习(CRL)确保驾驶安全,以手工设计的能量函数作为安全指标,定义安全与非安全区域。利用随状态变化的拉格朗日乘子,与安全策略同步学习,求解CRL问题。在高速场景实验中,本方法在安全性和性能上均优于传统MPC和SRL。
原文摘要 · Abstract (English)
Model predictive control (MPC) is widely used for motion planning, particularly in autonomous driving. Real-time capability of the planner requires utilizing convex approximation of optimal control problems (OCPs) for the planner. However, such approximations confine the solution to a subspace, which might not contain the global optimum. To address this, we propose using safe reinforcement learning (SRL) to obtain a new and safe reference trajectory within MPC. By employing a learning-based approach, the MPC can explore solutions beyond the close neighborhood of the previous one, potentially finding global optima. We incorporate constrained reinforcement learning (CRL) to ensure safety in automated driving, using a handcrafted energy function-based safety index as the constraint objective to model safe and unsafe regions. Our approach utilizes a state-dependent Lagrangian multiplier, learned concurrently with the safe policy, to solve the CRL problem. Through experimentation in a highway scenario, we demonstrate the superiority of our approach over both MPC and SRL in terms of safety and performance measures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。