arXiv:2503.05819eess.SYcs.RO2025-03被引 5

提出无监督轨迹采样方法,提升复杂环境下路径规划的探索能力。

An Unsupervised C-Uniform Trajectory Sampler with Applications to Model Predictive Path Integral Control

  • 用神经网络实现无监督的均匀轨迹采样,避免离散化空间计算
  • 在长时序下保持轨迹分布均匀性,均匀度接近原方法
  • 适合高曲率路径规划场景,显著提升控制性能

基于采样的模型预测控制器通常从固定简单分布(如正态或均匀分布)中采样控制输入,导致轨迹样本集中在均值轨迹附近,限制了探索能力,降低在复杂环境中的可行解发现概率。现有方法尝试通过重塑轨迹分布或增加采样熵来提升多样性。本文先前提出C-Uniform轨迹生成概念,可计算控制输入概率以实现配置空间均匀采样。但该方法因计算复杂度难以扩展。为此,本文提出Neural C-Uniform,一种无需依赖离散配置空间的无监督轨迹采样器,有效缓解可扩展性问题。实验表明,Neural C-Uniform在长时序下仍能保持与原方法相当的均匀度。进一步,本文构建CU-MPPI,将Neural C-Uniform集成至现有MPPI框架。仿真与真实世界实验显示,在最优解具有高曲率的场景中,CU-MPPI性能显著提升。

原文摘要 · Abstract (English)

Sampling-based model predictive controllers generate trajectories by sampling control inputs from a fixed, simple distribution such as the normal or uniform distributions. This sampling method yields trajectory samples that are tightly clustered around a mean trajectory. This clustering behavior in turn, limits the exploration capability of the controller and reduces the likelihood of finding feasible solutions in complex environments. Recent work has attempted to address this problem by either reshaping the resulting trajectory distribution or increasing the sample entropy to enhance diversity and promote exploration. In our recent work, we introduced the concept of C-Uniform trajectory generation [1] which allows the computation of control input probabilities to generate trajectories that sample the configuration space uniformly. In this work, we first address the main limitation of this method: lack of scalability due to computational complexity. We introduce Neural C-Uniform, an unsupervised C-Uniform trajectory sampler that mitigates scalability issues by computing control input probabilities without relying on a discretized configuration space. Experiments show that Neural C-Uniform achieves a similar uniformity ratio to the original C-Uniform approach and generates trajectories over a longer time horizon while preserving uniformity. Next, we present CU-MPPI, which integrates Neural C-Uniform sampling into existing MPPI variants. We analyze the performance of CU-MPPI in simulation and real-world experiments. Our results indicate that in settings where the optimal solution has high curvature, CU-MPPI leads to drastic improvements in performance.

路径规划轨迹采样强化学习控制算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。