arXiv:2412.10628cs.RO2024-12被引 4

六足机器人学会爬楼梯、避障和钻桌底,仅用仿真训练就能在真实环境稳定执行。

Versatile Locomotion Skills for Hexapod Robots

  • 分两阶段训练:先用强化学习学高阶策略,再用监督学习压缩模型至只依赖摄像头和姿态信息
  • 三类任务成功率均超90%,且无需实时关节反馈,可部署在低成本硬件上
  • 适合对复杂地形适应性强的机器人系统研究者,尤其关注仿真到现实迁移的团队

六足机器人因其稳定性、紧凑性和轻量化,适用于杂乱环境中的任务。其多关节腿部与可变高度躯干使其具备攀爬楼梯、在家庭或阁楼环境中钻过障碍物的能力。在先前阁楼横梁攀爬工作的基础上,我们训练了一个配备深度相机和视觉惯性里程计(VIO)的六足机器人,完成三项任务:爬楼梯、避障以及从桌子下穿行。模型仅通过仿真数据训练,可在无需实时关节状态反馈的低成本硬件上部署。采用教师-学生框架分两阶段训练:第一阶段利用强化学习,访问高度图和关节反馈等特权信息;第二阶段通过监督学习,将模型压缩为仅依赖本体感知——即前视深度图像和由追踪式VIO相机获取的机器人位姿。通过操控特权信息、构建仿真地形并优化奖励函数,训练出在非理想物理环境下仍具鲁棒性的技能。实验证明了良好的仿真到现实迁移效果,三项任务在物理实验中均取得高成功率。

原文摘要 · Abstract (English)

Hexapod robots are potentially suitable for carrying out tasks in cluttered environments since they are stable, compact, and light weight. They also have multi-joint legs and variable height bodies that make them good candidates for tasks such as stairs climbing and squeezing under objects in a typical home environment or an attic. Expanding on our previous work on joist climbing in attics, we train a legged hexapod equipped with a depth camera and visual inertial odometry (VIO) to perform three tasks: climbing stairs, avoiding obstacles, and squeezing under obstacles such as a table. Our policies are trained with simulation data only and can be deployed on lowcost hardware not requiring real-time joint state feedback. We train our model in a teacher-student model with 2 phases: In phase 1, we use reinforcement learning with access to privileged information such as height maps and joint feedback. In phase 2, we use supervised learning to distill the model into one with access to only onboard observations, consisting of egocentric depth images and robot pose captured by a tracking VIO camera. By manipulating available privileged information, constructing simulation terrains, and refining reward functions during phase 1 training, we are able to train the robots with skills that are robust in non-ideal physical environments. We demonstrate successful sim-to-real transfer and achieve high success rates across all three tasks in physical experiments.

六足机器人仿真到现实强化学习自主导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。