用二次模型降低采样方差,提升控制策略效率
Variance-Reduced Model Predictive Path Integral via Quadratic Model Approximation
- 将目标函数分解为已知二次模型与残差项,聚焦关键区域采样
- 在少样本下收敛更快,非光滑系统中性能优于传统MPPI
- 兼容梯度、结构或随机平滑信息,适用性广
基于采样的控制器(如模型预测路径积分,MPPI)虽灵活但常面临高方差和低样本效率问题。本文提出一种融合先验模型的混合方差缩减MPPI框架。核心思想是将目标函数分解为已知近似模型与残差项;由于残差仅反映模型与真实目标的偏差,其量级和方差通常远小于原目标。尽管该原则适用于各类建模方式,我们证明采用二次近似可导出闭式、模型引导的先验,有效集中采样于信息丰富区域。该框架不依赖几何信息来源,二次模型可由精确导数、结构近似(如高斯或拟牛顿法)或无梯度随机平滑构造。在标准优化基准、非线性欠驱动倒立摆控制任务及具接触动力学的复杂操作问题上验证,该方法在少样本条件下实现更快收敛与更优性能,表明其能显著提升样本代价高昂场景下的采样控制实用性。
原文摘要 · Abstract (English)
Sampling-based controllers, such as Model Predictive Path Integral (MPPI) methods, offer substantial flexibility but often suffer from high variance and low sample efficiency. To address these challenges, we introduce a hybrid variance-reduced MPPI framework that integrates a prior model into the sampling process. Our key insight is to decompose the objective function into a known approximate model and a residual term. Since the residual captures only the discrepancy between the model and the objective, it typically exhibits a smaller magnitude and lower variance than the original objective. Although this principle applies to general modeling choices, we demonstrate that adopting a quadratic approximation enables the derivation of a closed-form, model-guided prior that effectively concentrates samples in informative regions. Crucially, the framework is agnostic to the source of geometric information, allowing the quadratic model to be constructed from exact derivatives, structural approximations (e.g., Gauss- or Quasi-Newton), or gradient-free randomized smoothing. We validate the approach on standard optimization benchmarks, a nonlinear, underactuated cart-pole control task, and a contact-rich manipulation problem with non-smooth dynamics. Across these domains, we achieve faster convergence and superior performance in low-sample regimes compared to standard MPPI. These results suggest that the method can make sample-based control strategies more practical in scenarios where obtaining samples is expensive or limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。