arXiv:2510.00272cs.RO2025-10被引 2

为强化学习路径积分控制添加概率安全层,自动规避碰撞与越界。

BC-MPPI: A Probabilistic Constraint Layer for Safe Model-Predictive Path-Integral Control

  • 用概率代理模型评估轨迹可行性,动态调整采样权重。
  • 在100至1500次采样下,满足指定违规概率约束且保持安全裕度。
  • 无需调参或丢弃样本,适合可验证的自主系统集成。

模型预测路径积分(MPPI)控制近年来成为高非线性机器人任务中一种快速、无梯度的模型预测控制替代方案,但无法保证约束满足。我们提出贝叶斯约束MPPI(BC-MPPI),一种轻量级安全层,为每个状态和输入约束附加一个概率代理模型。在每次重规划步骤中,该代理返回候选轨迹的可行概率;该联合概率用于缩放候选样本的权重,自动降低可能发生碰撞或越界的轨迹权重,推动采样分布向安全子集收敛;无需人工调参的惩罚项或显式样本剔除。我们通过1000次离线仿真训练代理模型,并在MuJoCo中部署于四旋翼无人机,应对静态与动态障碍物。在K ∈ [100, 1500]次采样下,BC-MPPI保持安全裕度并满足预设违规概率要求。由于代理模型是独立、版本可控的实体,且运行时安全评分仅为单一标量,该方法天然适配可认证自主系统的验证与验证流程。

原文摘要 · Abstract (English)

Model Predictive Path Integral (MPPI) control has recently emerged as a fast, gradient-free alternative to model-predictive control in highly non-linear robotic tasks, yet it offers no hard guarantees on constraint satisfaction. We introduce Bayesian-Constraints MPPI (BC-MPPI), a lightweight safety layer that attaches a probabilistic surrogate to every state and input constraint. At each re-planning step the surrogate returns the probability that a candidate trajectory is feasible; this joint probability scales the weight given to a candidate, automatically down-weighting rollouts likely to collide or exceed limits and pushing the sampling distribution toward the safe subset; no hand-tuned penalty costs or explicit sample rejection required. We train the surrogate from 1000 offline simulations and deploy the controller on a quadrotor in MuJoCo with both static and moving obstacles. Across K in [100,1500] rollouts BC-MPPI preserves safety margins while satisfying the prescribed probability of violation. Because the surrogate is a stand-alone, version-controlled artefact and the runtime safety score is a single scalar, the approach integrates naturally with verification-and-validation pipelines for certifiable autonomous systems.

机器人控制安全强化学习概率约束路径规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。