arXiv:2410.10646cs.ROcs.AI2024-10中稿 · ICRA被引 32

用真实人群数据训练机器人安全导航,4小时即可学会避让复杂人流。

DR-MPC: Deep Residual Model Predictive Control for Real-world Social Navigation

  • 结合模型预测控制与无模型强化学习,逐步优化导航策略。
  • 在仿真中表现优于传统强化学习方法,在真实场景中仅用4小时数据就成功避障。
  • 通过识别异常状态提前预警,确保训练初期安全,适合实际部署的机器人导航。

机器人如何在复杂人流中安全导航?尽管模拟环境中的深度强化学习(DRL)有一定前景,但多数现有方法依赖无法捕捉真实人类运动细节的仿真器。为此,我们提出深度残差模型预测控制(DR-MPC),使机器人能基于真实人群导航数据快速、安全地完成强化学习。该方法将模型预测控制(MPC)与无模型DRL融合,克服了DRL对大量数据的需求和初始行为不安全的问题。DR-MPC以基于MPC的路径跟踪为起点,逐步学习更高效的与人交互方式。为进一步加速学习,引入安全模块,用于识别分布外状态并引导机器人避开潜在碰撞。仿真结果表明,DR-MPC显著优于以往方法,包括传统DRL和残差DRL模型。硬件实验显示,该方法仅需少于4小时的真实训练数据,即可让机器人在多种拥挤场景中成功导航,错误极少。

原文摘要 · Abstract (English)

How can a robot safely navigate around people with complex motion patterns? Deep Reinforcement Learning (DRL) in simulation holds some promise, but much prior work relies on simulators that fail to capture the nuances of real human motion. Thus, we propose Deep Residual Model Predictive Control (DR-MPC) to enable robots to quickly and safely perform DRL from real-world crowd navigation data. By blending MPC with model-free DRL, DR-MPC overcomes the DRL challenges of large data requirements and unsafe initial behavior. DR-MPC is initialized with MPC-based path tracking, and gradually learns to interact more effectively with humans. To further accelerate learning, a safety component estimates out-of-distribution states to guide the robot away from likely collisions. In simulation, we show that DR-MPC substantially outperforms prior work, including traditional DRL and residual DRL models. Hardware experiments show our approach successfully enables a robot to navigate a variety of crowded situations with few errors using less than 4 hours of training data.

机器人导航强化学习真实世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。