arXiv:2511.02015cs.RO2025-11

用斯坦因方法优化采样分布,提升模型预测路径积分控制性能

Stein-based Optimization of Sampling Distributions in Model Predictive Path Integral Control

  • 结合斯坦因变分梯度下降动态调整动作采样分布
  • 在低粒子数下仍实现比现有方法更优的系统表现
  • 适合需要高效实时控制的机器人系统

本文提出一种基于斯坦因变分梯度下降(SVGD)优化采样分布的模型预测路径积分(MPPI)控制方法。传统MPPI假设动作分布为单峰高斯,易因样本匮乏导致预测偏差,且在可微仿真中对成本梯度噪声敏感。通过在环境步之间引入SVGD更新,本文构建了斯坦因优化路径积分推断(SOPPI)算法,可在运行时动态调整噪声分布,更好捕捉动作采样特性,同时计算开销可控。在平面倒立摆、7自由度机械臂和平面双足行走者上验证,SOPPI在多种超参数设置下均优于当前最优MPPI算法,并在较低粒子数下仍保持有效性。

原文摘要 · Abstract (English)

This paper introduces a method for Model Predictive Path Integral (MPPI) control that optimizes sample generation towards an optimal trajectory through Stein Variational Gradient Descent (SVGD). MPPI relies upon predictive rollout of trajectories sampled from a distribution of possible actions. Traditionally, these action distributions are assumed to be unimodal and represented as Gaussian. The result can lead suboptimal rollout predictions due to sample deprivation and, in the case of differentiable simulation, sensitivity to noise in the cost gradients. Through introducing SVGD updates in between MPPI environment steps, we present Stein-Optimized Path-Integral Inference (SOPPI), an MPPI/SVGD algorithm that can dynamically update noise distributions at runtime to better capture action sampling distributions without an excessive increase in computational requirements. We demonstrate the efficacy of SOPPI through experiments on a planar cart-pole, 7-DOF robot arm, and a planar bipedal walker. These results indicate improved system performance compared to state-of-the-art MPPI algorithms across a range of hyper-parameters and demonstrate feasibility at lower particle counts.

强化学习路径规划优化算法机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。