提出混合优化机制,让机器人在复杂障碍中更稳定地规划路径。
Beyond Pure Sampling: Hybrid Optimization Mechanisms for Non-Convex Model Predictive Control

- 先用梯度优化,再通过采样打破局部最优陷阱。
- 高维系统下成功率高于传统方法,且更稳定。
- 适合需要可靠路径规划的无人机等真实机器人系统。
本文研究基于最大熵微分动态规划(ME-DDP)框架下非凸模型预测控制(MPC)的优化机制。针对由非线性动力学、多重障碍等引起的非凸代价景观,梯度法常陷入次优局部极小值。提出双阶段优化机制:第一阶段利用DDP探索代价景观梯度,第二阶段通过从动作价值函数逆海森矩阵定义的策略中采样来破坏优化过程。分析了三种ME-DDP变体:单模高斯ME-DDP、多模高斯ME-DDP和斯坦变换变分DDP。在四种机器人系统于杂乱环境中的导航任务中,对三种ME-DDP变体与确定性DDP及最成功的采样方法之一——模型预测路径积分(MPPI)控制进行广泛对比,采用三种对应于ME-DDP的策略参数化与更新规则。结果表明,在低维系统中,该框架持续优于MPPI;在高维系统中,MPPI偶可发现激进动作实现更快推进,但本方法保持更高且更稳定的成功概率。最后通过四旋翼在密集非凸障碍场的硬件实验,验证了该框架在真实部署中的有效性与鲁棒性。
原文摘要 · Abstract (English)
This paper investigates the optimization mechanisms of non-convex Model Predictive Control (MPC) using the Maximum Entropy Differential Dynamic Programming (ME-DDP) framework. Navigating non-convex cost landscapes induced by nonlinear dynamics, multiple obstacles, etc. remains a fundamental challenge in robotics, where gradient-based methods frequently converge to suboptimal local minima. We demonstrate a dual-step optimization mechanism designed to overcome these traps. (1) an initial phase of using DDP to exploit the gradient of the cost landscape, followed by (2) disruption of the optimization via sampling from policies characterized by the inverse Hessian of the action-value function. We provide a rigorous analysis of this sampling mechanism of three ME-DDP variants: Unimodal Gaussian ME-DDP, Multimodal Gaussian ME-DDP, and Stein Variational DDP. Furthermore, with navigation tasks of four robotic systems under cluttered environments, we conduct extensive benchmarking of three variants of the ME-DDP, against deterministic DDP, and one of the most successful sampling-based schemes, Model Predictive Path Integral (MPPI) control with three policy parameterizations and update laws that correspond to those of ME-DDPs. The results show that in low-dimensional systems where the cost landscapes are relatively simple and local information is sufficiently representative, our framework consistently outperforms MPPIs. In high-dimensional systems, MPPI can occasionally discover aggressive maneuvers that enable it to steer the systems faster than DDP-based methods, whereas our method maintains a higher, more stable success rate. Finally, we validate the practical efficacy of the framework through hardware experiments with a quadrotor navigating a dense, non-convex obstacle field, confirming the robustness of the proposed framework for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。