arXiv:2606.24991eess.SYcs.LG2026-06中稿 · IFAC World Congres…

将未来信息融入MPC,实现带前瞻的最优决策

Solving Markov Decision Processes with Future Information via MPC

  • 用参数化MPC建模含未来信息的MDP最优策略
  • 理论证明该方法可精确表示最优价值函数与策略
  • 适合需实时规划且有未来预测的机器人任务

模型预测控制(MPC)广泛应用于工业和机器人系统中,通过有限时域优化规划来实现约束满足和领域知识嵌入。然而,传统MPC通常无法为马尔可夫决策过程(MDP)提供最优策略。近期结合强化学习(RL)的MPC方法通过将MPC视为MDP最优策略的参数化模型,并利用数据调整其参数,缓解了这一问题。但这些方法大多针对经典MDP,而许多现实问题在决策时刻包含未来信息(如预测、价格或参考轨迹),必须作为状态的一部分才能实现最优决策。现有MPC-RL方法未直接处理这种扩展状态结构,因此亟需解决如何将未来信息融入MPC以获得最优策略的问题。本文建立了参数化MPC能够精确表示含未来信息的MDP最优值函数与策略的结构条件,并证明该参数化MPC可作为结构化函数逼近器,其参数可通过强化学习进行学习。该方法在带有未来参考信息的点质量竞速任务上进行了验证。

原文摘要 · Abstract (English)

Model Predictive Control (MPC) is widely used in industrial and robotic systems for enforcing constraints and embedding domain knowledge through finite-horizon optimization-based planning. However, despite these strengths, an MPC scheme typically does not yield optimal policies for sequential decision-making problems formulated as Markov Decision Processes (MDPs). Recent combinations of MPC with Reinforcement Learning (RL) alleviate this issue by treating MPC as a parameterized model of the optimal policy of an MDP and adjusting its parameters using data. While these approaches typically consider classical MDPs, many real-world problems include future information--such as forecasts, prices, or reference trajectories--at decision time, which must be included in the MDP state for optimal decision-making. Current MPC-RL approaches do not directly account for this augmented-state structure, raising the question of how to incorporate future information into MPC to obtain an optimal policy. This work establishes the structural requirements under which a parameterized MPC can exactly represent the optimal value functions and policy of an MDP with future information. We further demonstrate that such a parameterized MPC can serve as a structured function approximator, with its parameters learned using RL. The approach is illustrated on a point-mass racing task with future reference information.

MPC强化学习最优控制未来信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。