让机器人在参数未知时仍能安全学习并改进性能
Parameter-Robust MPPI for Safe Online Learning of Unknown Parameters
- 用粒子信念+在线学习实时更新不确定参数
- 同时优化性能与安全备份轨迹,成功率更高
- 适合动态环境中需持续安全运行的机器人
部署于动态环境中的机器人必须在关键物理参数不确定或随时间变化时仍保持安全。我们提出参数鲁棒模型预测路径积分(PRMPPI)控制框架,将在线参数学习与概率安全约束相结合。PRMPPI通过斯坦纳变分梯度下降维护参数的粒子信念,利用置信预测评估安全约束,并并行优化一个以性能为导向的基准轨迹和一个以安全为重点的备份轨迹。该控制器初始阶段保持谨慎,随着参数学习逐步提升性能,全程确保安全性。仿真与硬件实验表明,相较于基线方法,PRMPPI具有更高的成功率、更低的跟踪误差和更精确的参数估计。
原文摘要 · Abstract (English)
Robots deployed in dynamic environments must remain safe even when key physical parameters are uncertain or change over time. We propose Parameter-Robust Model Predictive Path Integral (PRMPPI) control, a framework that integrates online parameter learning with probabilistic safety constraints. PRMPPI maintains a particle-based belief over parameters via Stein Variational Gradient Descent, evaluates safety constraints using Conformal Prediction, and optimizes both a nominal performance-driven and a safety-focused backup trajectory in parallel. This yields a controller that is cautious at first, improves performance as parameters are learned, and ensures safety throughout. Simulation and hardware experiments demonstrate higher success rates, lower tracking error, and more accurate parameter estimates than baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。