arXiv:2502.02133eess.SYcs.AI2025-02综述被引 40

融合模型预测控制与强化学习,提升系统控制性能。

Synthesis of Model Predictive Control and Reinforcement Learning: Survey and Classification

  • 以演员-评论家强化学习为基础,结合模型预测控制的在线优化能力。
  • 通过融合两类方法,实现更优的闭环控制性能。
  • 适合对智能控制、机器人决策感兴趣的读者。

模型预测控制(MPC)和强化学习(RL)是马尔可夫决策过程中的两种成功控制技术,均源于相似的基本原理,并广泛应用于机器人、过程控制、能源系统和自动驾驶等领域。尽管二者具有相似性,但它们分别源自不同社区,满足不同需求,遵循不同的范式。关键技术差异,尤其是环境模型在算法中的作用,导致两类方法优势互补。由于其正交优势,近年来结合两者的研究兴趣显著增加,涌现出大量复杂且不断发展的思想。本文阐明了两者在差异、相似性和基本原理上的联系,为现有工作提供了分类框架。特别地,以通用的演员-评论家强化学习方法为基准,探讨如何利用MPC的在线优化特性提升策略的整体闭环性能。

原文摘要 · Abstract (English)

The fields of MPC and RL consider two successful control techniques for Markov decision processes. Both approaches are derived from similar fundamental principles, and both are widely used in practical applications, including robotics, process control, energy systems, and autonomous driving. Despite their similarities, MPC and RL follow distinct paradigms that emerged from diverse communities and different requirements. Various technical discrepancies, particularly the role of an environment model as part of the algorithm, lead to methodologies with nearly complementary advantages. Due to their orthogonal benefits, research interest in combination methods has recently increased significantly, leading to a large and growing set of complex ideas leveraging MPC and RL. This work illuminates the differences, similarities, and fundamentals that allow for different combination algorithms and categorizes existing work accordingly. Particularly, we focus on the versatile actor-critic RL approach as a basis for our categorization and examine how the online optimization approach of MPC can be used to improve the overall closed-loop performance of a policy.

控制理论强化学习MPC融合方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。