将DRL与MPC分层协作,提升多类型交通网控制效率与稳定性。
Sharing the Control Authority Between Deep Reinforcement Learning and Model Predictive Control: Application to Multi-Class Transportation Networks

- DRL处理高频控制,MPC负责低频优化,分权协作降低计算负担。
- 在模型不匹配和噪声需求下,约束执行效果优于对比方法。
- 适合需实时响应且模型不精确的复杂交通网络场景。
多类型交通网络(即混合车型网络)是复杂的系统,难以控制。近年来,深度强化学习(DRL)通过与环境交互学习控制策略,模型预测控制(MPC)则利用系统模型优化控制输入,被越来越多地应用于交通网络控制。然而,大规模网络中非线性动态和高维状态空间限制了DRL在时间受限训练下的学习能力,并增加MPC的计算耗时,阻碍了在有限算力下的实时部署。此外,MPC依赖精确的网络模型,而复杂系统如多类型交通网络往往缺乏准确模型。本文提出一种新型的DRL-MPC框架,将控制权在DRL与MPC间分配,结合DRL快速在线计算与无模型特性,以及MPC内置优化与约束处理能力。在分层结构中,MPC位于高层,生成更新频率较低的控制输入,适应其高计算开销;DRL位于底层,以快速在线部署方式生成高频控制输入。该框架在多类型高速公路网络上进行评估,对比了分层MPC控制器与混合状态反馈-MPC控制器,涵盖模型失配与噪声交通需求场景。结果表明,所提框架性能优于混合状态反馈-MPC控制器,相比分层MPC显著降低在线计算时间,并在模型失配下提供更强的约束保障能力。
原文摘要 · Abstract (English)
Transportation networks, in particular multi-class transportation networks (i.e., networks with mixed vehicle types), are complex systems that are challenging to control. Recently, Deep Reinforcement Learning (DRL), which learns control policies from interactions with the environment, and Model Predictive Control (MPC), which uses a system model to optimize control inputs, have been increasingly utilized for transportation network control. However, nonlinear system dynamics and high-dimensional state spaces in large-scale networks limit DRL's learning capacity under time-constrained training and increase MPC's computation time, hindering real-time implementation with limited computational resources. Moreover, MPC depends on an accurate network model, which is often unavailable for complex systems such as multi-class transportation networks. This paper proposes a novel DRL-MPC framework for multi-class transportation networks that divides control authority between DRL and MPC, combining DRL's fast online computation and model independence with MPC's built-in optimization and constraint-handling capabilities. In the hierarchical framework, MPC operates at the higher level and determines low-frequency control inputs whose slower update rate accommodates its high computation time, while DRL operates at the lower level and determines high-frequency control inputs using its fast online deployment. The framework is evaluated on a multi-class freeway network against a hierarchical MPC controller and a hybrid state-feedback-MPC controller, including scenarios with model mismatch and noisy traffic demands. Results show that the proposed framework outperforms the hybrid state-feedback-MPC controller, substantially reduces online computation time compared with the hierarchical MPC controller, and provides more effective constraint enforcement under model mismatch.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。