arXiv:2409.02383cs.RO2024-09被引 16

用强化学习让轮式机器人学会攀爬陡坡巨石。

Reinforcement Learning for Wheeled Mobility on Vertically Challenging Terrain

  • 通过仿真试错训练,端到端学习轮地交互策略。
  • 在模拟和真实机器人上均实现垂直地形稳定通过。
  • 无需复杂建模,适合普通轮式机器人升级使用。

在包含陡坡和崎岖巨石的垂直挑战性地形上进行非公路导航,对轮式机器人在路径规划层面实现平滑无碰撞轨迹,以及在控制层面避免翻倒或卡住,都提出了巨大挑战。考虑到轮-地形相互作用的复杂模型,我们基于自研的Chrono多物理引擎仿真器,开发了一个端到端的强化学习(RL)系统,使自主车辆通过模拟中的试错经验学习轮式移动能力。该方法采用近端策略优化(PPO)并结合地形难度课程,基于奖励函数优化策略,鼓励向目标前进并惩罚过大的滚动与俯仰角度,从而避免了复杂且昂贵的运动学动力学建模、规划与控制需求。此外,我们在仿真环境中展示了实验结果,并将该方法部署于真实的Verti-4-Wheeler(V4W)平台上,验证了强化学习可为传统轮式机器人赋予此前难以实现的垂直地形通行能力。

原文摘要 · Abstract (English)

Off-road navigation on vertically challenging terrain, involving steep slopes and rugged boulders, presents significant challenges for wheeled robots both at the planning level to achieve smooth collision-free trajectories and at the control level to avoid rolling over or getting stuck. Considering the complex model of wheel-terrain interactions, we develop an end-to-end Reinforcement Learning (RL) system for an autonomous vehicle to learn wheeled mobility through simulated trial-and-error experiences. Using a custom-designed simulator built on the Chrono multi-physics engine, our approach leverages Proximal Policy Optimization (PPO) and a terrain difficulty curriculum to refine a policy based on a reward function to encourage progress towards the goal and penalize excessive roll and pitch angles, which circumvents the need of complex and expensive kinodynamic modeling, planning, and control. Additionally, we present experimental results in the simulator and deploy our approach on a physical Verti-4-Wheeler (V4W) platform, demonstrating that RL can equip conventional wheeled robots with previously unrealized potential of navigating vertically challenging terrain.

强化学习机器人越野导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。