为离散LTI系统中的MPPI控制提供了有限样本闭环稳定性保证。
Finite-Sample Closed-Loop Stability of Model Predictive Path Integral Control for Linear Time-Invariant Systems
- 将MPPI视为LQR的随机扰动,利用蒙特卡洛采样逼近反馈控制。
- 在有限样本下实现期望状态指数衰减,受过程噪声、近似误差和采样失败三重限制。
- 给出可计算的最小采样数阈值,适用于工程中对稳定性的严格要求场景。
针对带有加性高斯过程扰动的离散时间线性时不变(LTI)系统,本文建立了模型预测路径积分(MPPI)控制的有限样本闭环稳定性保证。关键观察发现:对于无约束的LTI/二次系统,采用DARE终端代价时,有限时域MPC策略的首步控制律与无限时域LQR策略完全一致。因此,有限样本下的MPPI可被分析为对LQR的随机扰动。首先,证明了MPPI控制律以高概率逼近LQR反馈,其误差由随采样数减少的蒙特卡洛项和在有限温度下持续存在的无限样本温度偏差组成;相关常数显式依赖于时域相关的堆叠代价矩阵,表明证书由所选规划时域参数化。其次,通过李雅普诺夫扰动论证,证明了期望下的实用指数稳定性。在有限运行时域内保持于紧致李雅普诺夫子集的样本路径上,期望状态范数呈指数衰减,存在三个残差层:过程噪声层、MPPI近似层和因每步采样失败概率带来的置信层。所需最小采样阈值可由DARE解、LQR稳定性裕度、MPPI采样参数、温度及规划时域显式计算得出。在无限样本与温度偏差消失的联合极限下,结果恢复经典的随机LQR稳定性界。
原文摘要 · Abstract (English)
We establish finite-sample closed-loop stability guarantees for Model Predictive Path Integral (MPPI) control applied to discrete-time Linear Time-Invariant (LTI) systems with additive Gaussian process disturbances. The key observation is that, for unconstrained LTI/quadratic systems with the DARE terminal cost, the exact finite-horizon MPC law has the same first control action as the infinite-horizon LQR law for every planning horizon. Thus, finite-sample MPPI can be analyzed as a stochastic perturbation of LQR. First, we show that the MPPI control law approximates the LQR feedback with high probability. The approximation error decomposes into a Monte Carlo term that decreases with the sample count and an infinite-sample temperature bias that persists at finite temperature but vanishes as the temperature is reduced. The resulting constants are written in terms of the horizon-dependent stacked cost matrices, making explicit that the finite-sample certificate is parametrized by the selected planning horizon. Second, we use a Lyapunov perturbation argument to prove practical exponential stability in expectation. On sample paths that remain in a compact Lyapunov sublevel set over a finite operating horizon, the expected state norm decays exponentially up to three residual floors: a process-noise floor, an MPPI approximation floor, and a confidence floor from the per-step sampling failure probability. The sufficient sample threshold is explicit and computable from the DARE solution, LQR stability margin, MPPI sampling parameters, temperature, and planning horizon. In the joint limit of infinite samples and vanishing temperature bias, the result recovers the stochastic LQR stability bound.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。