用激光雷达隐状态建模,让机器人在密集人群里更安全避障。
RAY-TOLD: Ray-Based Latent Dynamics for Dense Dynamic Obstacle Avoidance with TDMPC

- 基于激光雷达的隐状态模型压缩传感器数据,学习长期避障策略。
- 混合物理规划与学习策略,碰撞率显著降低。
- 适合高密度动态障碍场景下的移动机器人导航应用。
密集动态人群对自主移动机器人构成持续挑战。纯反应式规划方法(如模型预测路径积分控制)因预测时长有限,常在复杂场景中陷入局部最优。为此,我们提出基于射线的任务导向隐状态动力学模型(RAY-TOLD),融合障碍物信息到隐状态动力学中,并结合基于物理的MPPI鲁棒性与强化学习的长时前瞻能力。该方法采用以激光雷达为中心的隐状态动力学模型,将高维传感器数据压缩为紧凑状态表示,从而学习终端价值函数和策略先验。引入策略混合采样策略,在MPPI候选轨迹中加入由学习策略生成的路径,有效引导规划朝向目标,同时保持运动可行性。在具有高密度动态障碍物的随机环境中的大量测试表明,本方法优于基线MPPI,显著降低碰撞率。结果验证了将短时物理模拟与学习得到的长时意图相结合,可大幅提升导航可靠性与安全性。
原文摘要 · Abstract (English)
Dense, dynamic crowds pose a persistent challenge for autonomous mobile robots. Purely reactive planning methods, such as Model Predictive Path Integral (MPPI) control, often fail to escape local minima in complex scenarios due to their limited prediction horizon. To bridge this gap, we propose Ray-based Task-Oriented Latent Dynamics (RAY-TOLD), a hybrid control architecture that integrates obstacle information into latent dynamics and utilizes the robustness of physics-based MPPI with the long-horizon foresight of reinforcement learning. RAY-TOLD leverages a LiDAR-centric latent dynamics model to encode high-dimensional sensor data into a compact state representation, enabling the learning of a terminal value function and a policy prior. We introduce a policy mixture sampling strategy that augments the MPPI candidate population with trajectories derived from the learned policy, effectively guiding the planner towards the goal while maintaining kinematic feasibility. Extensive tests in a stochastic environment with high-density dynamic obstacles demonstrate that our method outperforms the MPPI baseline, reducing the collision rate. The results confirm that blending short-horizon physics-based rollouts with learned long-horizon intent significantly enhances navigation reliability and safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。